What happened
- A paper announced on September 29 evaluates whether switching model versions improves probability estimates on prediction-market questions.
- The benchmark gathers 3,000 already-resolved binary questions across nine domains. Six Claude and Qwen variants were tested with a zero-shot protocol and only the question statement.
- On the segment after the training cutoff, the four Claude models score between 0.183 and 0.192 Brier, improving between 0.024 and 0.033 over a base-rate predictor. Qwen 32B does not significantly beat that reference.
- The conclusion that organizes the paper: differences between event categories within a single model are larger than differences between Claude variants. Moving up a version or a tier produced no significant improvements under this protocol.
Why it matters
- The comparison against a base rate is what almost never appears in launch announcements. Without it, an improvement of 0.03 can be read as a leap or as noise, and no one can tell which.
- For a team using these models to estimate demand, the risk of a supplier or the odds that a campaign works, the result is direct: paying for the top tier isn’t worth it for this task.
- Differences between domains matter more than differences between models. It is better to measure in your own domain than to inherit a choice from a published average.
The number
0.024 to 0.033 improvement over the base rate, with no significant differences between versions.
Context
Anthropic has just released Sonnet 5.5, with a claimed jump from 10.3% to 70.6% on a terminal test, and earlier cut the cost of its largest model by 40% in a price cut that shifted cybersecurity to another model. Coding tests move sharply between versions. Probability tests, according to this paper, do not.
What’s next
- No announced date for releasing the 3,000-question resolved benchmark.
- The authors did not evaluate models from other frontier families. No timeline to extend it.
Bottom line
Xiaomi released open weights that beat Opus 5 on two tests, and the comparison was discussed as if one test spoke for all. This paper shows the opposite from the other side: there are tasks where changing models, in either direction, doesn’t move the result.
Sources
- Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets, arXiv, announced September 29, 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.


