Agent evaluation

Email agents execute 99.7% of their calls and solve 33.5% of tasks


A benchmark of 206 office scenarios separates two things that get confused fairly often: the tool responding and the task getting done.

September 29, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

33.5% versus 99.7%: scenarios passed against tool calls executed without failure.

Context

We had already written about agents that declare themselves done 38 points more often than the evaluator approves and about the judgment measurement where the top-scoring system gets 59.7% right. EmailBench adds the most common case in an office, which is also the least measured.

What’s next

Bottom line

Microsoft rebuilt Copilot around an agent that doesn’t switch off and put a taximeter on the work. A taximeter that counts executed calls charges just the same for the tasks left half done. That is the gap this benchmark has just put into numbers.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.

Related notes

← All notes