How's it going, agents?

Microsoft measured agents' judgment, and the best one gets it right 59.7% of the time


The study evaluates decisions with no single answer in long tasks and concludes that expanding the reasoning budget doesn't improve the result.

September 23, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

59.7%. The accuracy of the best model evaluated on decisions where there’s no single answer.

Context

It fits with what we’d already been noting about AI engines that disagree with each other when describing the same brand and about how machine-made content gets treated. The problem isn’t the accuracy of the facts; it’s choosing between defensible alternatives.

What’s next

Bottom line

The industry has spent two years debating whether models hallucinate. This measurement asks something else: if they get the facts right but choose badly, the text still comes out wrong.

Sources

Edited by Rodrigo Cornejo. How we select, verify and correct each note.

Related notes

← All notes