What happened
- A team published EmailBench on September 29, a benchmark of 206 email and productivity scenarios spread across 16 task categories.
- The environment is closed and synthetic, inspired by the Enron corpus, with a typed email interface specification and provider-neutral names. The evaluation mixes 258 executable assertions with 211 rubrics judged by a model.
- Eight model configurations were tested on the same single-user corpus. The top-scoring one passes 33.5% of the scenarios.
- The contrast is the finding: that same configuration completed 99.7% of its tool calls with no observed interface failure. Pass rates vary widely between task categories.
Why it matters
- Almost all the agent dashboards sold today measure the first thing: successful calls, response time, interface errors. That dashboard would have shown 99.7% while two out of three tasks went undone.
- For a team in Chile weighing whether to automate its inbox, the useful question stops being whether the agent connects well to email. It is whether someone is checking the result against what was asked, and how often.
- Corporate email tasks mix information retrieval, state changes, reasoning about dates and multi-step coordination. The benchmark shows the difficulty is not evenly spread, so automating by category pays off more than automating the whole inbox.
The number
33.5% versus 99.7%: scenarios passed against tool calls executed without failure.
Context
We had already written about agents that declare themselves done 38 points more often than the evaluator approves and about the judgment measurement where the top-scoring system gets 59.7% right. EmailBench adds the most common case in an office, which is also the least measured.
What’s next
- The authors leave as pending work broadening tool coverage, testing with several people and repeating runs. No announced date.
- No timeline for releasing the corpus or the evaluation environment.
Bottom line
Microsoft rebuilt Copilot around an agent that doesn’t switch off and put a taximeter on the work. A taximeter that counts executed calls charges just the same for the tasks left half done. That is the gap this benchmark has just put into numbers.
Sources
- EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks, arXiv, announced September 29, 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.


