What happened
- Rahul Khedar, Mayank Malhotra and Avinash Karn describe Augur, a system that rehearses offline the reaction to a product or policy change before it goes out. They submitted it to arXiv on September 24, 2026.
- The procedure builds a knowledge graph from the change’s documents, populates a market of simulated people, runs the interaction and returns an auditable decision memo recommending one of five actions.
- To measure it they built Gold-50: fifty real product and policy episodes whose outcome is already known, checked against the public record. On that basis they score the five-way launch verdict.
- Blind judges from four model families concluded that the synthetic reaction recovers between 67% and 90% of the concerns the public actually raised, with the biggest contribution where the decision was hardest.
Why it matters
- The paper’s central finding is uncomfortable for anyone buying tools: most of the measured difference between large commercial models and open-weights models fine-tuned in-house didn’t come from capability; it came from a poorly specified evaluation. With the same model, the same cases and the same evaluator, one system went from 0% to 73% just because of how the prompt wrapper was written.
- Defining the decision taxonomy inside the prompt, without changing the model, raised every large model by between 24 and 34 percentage points. Anyone evaluating AI vendors with a poorly defined test of their own will buy based on a difference that doesn’t exist.
- The authors also warn that the full pipeline amplifies a systematic pessimism bias. A system that exaggerates a campaign’s risk is as costly as one that ignores it; it’s just that the error isn’t visible.
The number
From 0% to 73% accuracy with the same weights, the same cases and the same evaluator.
Context
The inconsistency of answers depending on how the question is asked had already been measured for brands: two AI engines name the same brand in 81% and 43% of their answers. And the question of judgment, not capability, is the same one opened by Microsoft’s test where the best model gets 59.7% right.
What’s next
- No timelines announced. The authors say the pipeline that regenerates every number and figure in the paper is available on request, with no public repository or publication date.
Bottom line
The paper measures a tool for rehearsing launches and ends up demonstrating something more useful: that much of what we call a model difference is a difference in wording.
Sources
- Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes, arXiv 2609.29952, September 24, 2026.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


