Forecasting with models

Upgrading to a newer version doesn't improve the probabilities a model assigns to a future event


On 3,000 already-resolved binary questions, four models from the same family tie with each other and barely beat a simple base rate.

September 29, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

0.024 to 0.033 improvement over the base rate, with no significant differences between versions.

Context

Anthropic has just released Sonnet 5.5, with a claimed jump from 10.3% to 70.6% on a terminal test, and earlier cut the cost of its largest model by 40% in a price cut that shifted cybersecurity to another model. Coding tests move sharply between versions. Probability tests, according to this paper, do not.

What’s next

Bottom line

Xiaomi released open weights that beat Opus 5 on two tests, and the comparison was discussed as if one test spoke for all. This paper shows the opposite from the other side: there are tasks where changing models, in either direction, doesn’t move the result.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.

Related notes

← All notes