Agent evaluation

AI agents say they're done up to 38 points more often than the grader approves


Seven models met between 79.6% and 86.4% of 509 instructions, and reported work finished up to 37.9 points above the official evaluation.

September 26, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

Between 28.7 and 37.9 percentage points separate the agent’s completion notice from the grader’s approval.

Context

Measuring agents’ judgment on long tasks had already left Microsoft’s best model at 59.7% accuracy. And the question of who averages and by what rule came up in the European panel where declared risk falls by up to 37 points depending on the aggregation.

What’s next

Bottom line

The specification existed in every case measured: it was in the prompt. What didn’t exist was someone other than the agent authorized to say it had been met.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes