Notes

A model's declared risk drops by up to 37 points depending on how it's averaged


A public dashboard organizes 19 tests under the European code's four categories. Switching from the average to the worst case sends scores plummeting.

September 24, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

A drop of between 14 and 37 points across 18 models when using worst-case aggregation instead of the average.

Context

The European framework is the one that has gone furthest in demanding evidence, although its calendar moved: the bloc postponed the application of the high-risk rules. In parallel, the labs published their own criteria for external evaluations, as when OpenAI released its priorities for third-party assessments without naming any evaluator. This work is the open counterpart.

What’s next

Bottom line

The debate over model safety has been fought over the published numbers. This work moves the debate one step back, to who chooses the formula those numbers are calculated with, which until now has been the same party that publishes them.

Sources

Editor’s note (Rodrigo Cornejo): I didn’t understand much. Are we safe?

Related notes

← All notes