What happened
- On September 23, a team published on arXiv an open data pipeline and an interactive dashboard for gathering evidence of systemic risk under the European Union’s code of practice for general-purpose models.
- The dashboard organizes 19 public tests into the four categories that code sets out: chemical, biological, radiological and nuclear threats; offensive cyber; harmful manipulation; and loss of control.
- The finding behind the work: when aggregation switches from averaging to worst case across the 18 models evaluated, scores fall by between 14 and 37 points.
- The authors validated the method with automated judges that reach human-level agreement (κ between 0.78 and 0.82) and with an independent audit that confirmed 83% of their perturbations preserved the original harm. A test with 21 users showed greater perceived transparency.
Why it matters
- The result doesn’t say the models are more dangerous than declared. It says the safety figure depends on a methodological decision that’s almost never made explicit. A difference of 14 to 37 points is a range where any conclusion fits.
- For anyone buying AI software in Chile or the region, this changes the question to ask the vendor. It’s not whether the model passed the evaluation; it’s how the results were aggregated. An average hides the worst case, and the worst case is the one that reaches production on some ordinary Tuesday.
- And there’s a consequence for the audit ecosystem being built: if an open dashboard with 19 public tests can move scores that much, the evaluations companies themselves report without detailing their aggregation are worth less than they seem.
The number
A drop of between 14 and 37 points across 18 models when using worst-case aggregation instead of the average.
Context
The European framework is the one that has gone furthest in demanding evidence, although its calendar moved: the bloc postponed the application of the high-risk rules. In parallel, the labs published their own criteria for external evaluations, as when OpenAI released its priorities for third-party assessments without naming any evaluator. This work is the open counterpart.
What’s next
- Preprint without peer review. The dashboard is open source.
- The authors don’t announce coverage of additional models or an update frequency.
- No timelines announced.
Bottom line
The debate over model safety has been fought over the published numbers. This work moves the debate one step back, to who chooses the formula those numbers are calculated with, which until now has been the same party that publishes them.
Sources
- An Open Pipeline and Dashboard for Systemic-Risk Evidence, arXiv 2609.28335, September 23, 2026.
Editor’s note (Rodrigo Cornejo): I didn’t understand much. Are we safe?


