What happened
- On September 24, a Princeton team posted a study to arXiv on what happens when some agents in a deliberating group don’t act in good faith.
- The main conclusion, according to the text: what determines vulnerability isn’t the size of the group but the proportion of deceptive agents. The defection rate — how many initially correct agents end up giving a wrong answer — grows linearly with that proportion.
- The comparison the authors draw with classic studies of human conformity is the hard fact: people reliably give in only when the confederates inducing the error are a majority. Agents defect regularly even when the deceivers remain a minority.
- Susceptibility depends on which models are talking to each other, and above all on which model is on the honest side.
- One result the authors flag as unexpected: letting the deceivers coordinate in private makes them less effective, not more.
Why it matters
- The commercial argument for multi-agent systems is that several instances check each other and consensus corrects individual error. This paper measures the opposite in the presence of a disloyal actor: adding agents isn’t a defense, because the adversary scales with the group.
- For teams in Chile building workflows with several chained agents — one drafts, another reviews, another approves — the operational question changes. It isn’t how many reviewers to add, it’s which model holds the honest seat and whether any link can receive instructions from outside.
- There’s a consequence the paper doesn’t develop: in a real chain, the disloyal agent doesn’t need to be malicious. One badly configured agent, with a contaminated system message or reading a manipulated page, is enough to fill that role without anyone having decided it.
The number
Defection rises linearly with the proportion of deceptive agents, regardless of how many agents the group has.
Context
Our note on agents that turn autonomy into risk described the problem in a single agent. This study carries it over to the arrangement the industry proposes as the solution.
What’s next
- No timelines announced for releasing the code or extending the evaluation to more model families.
- The paper is at version 1 and does not declare peer review.
Bottom line
Measuring an agent’s judgment on long tasks was already hard: Taste-Bench left the best model at 59.7%. Measuring it while another agent is pushing it in the wrong direction doesn’t have a standard test yet.
Sources
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


