What happened
- On September 22, a team including Microsoft researchers published “The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks,” a 33-page paper, on arXiv.
- The work defines judgment (taste, in the original) as the ability to decide well at the points of a long task where there’s no single correct answer, and proposes Taste-Bench to measure it.
- The benchmark builds itself: it extracts decision forks from the trajectories of agents that attempted the same task in parallel, with no human annotation.
- Across frontier models, the best one answers 59.7% of the questions correctly. Expanding the reasoning budget didn’t improve the result. Distilling from a model with better judgment did, even on new tasks and on SWE-bench Pro.
Why it matters
- The entire apparatus for evaluating agents measures whether the result is correct. This one measures whether the decision was good when several were defensible, which is exactly the terrain where branding, copywriting and design work is decided.
- For any team delegating pieces of communication to an agent, the operational finding is uncomfortable: giving it more time to think doesn’t fix judgment calls. What moves them is the example of someone who already decides well.
- A consequence the paper doesn’t spell out: if judgment is transmitted through distillation, anyone with a large archive of well-made decisions of their own has an asset that can’t be bought by subscription. And anyone without one inherits the internet’s average judgment.
The number
59.7%. The accuracy of the best model evaluated on decisions where there’s no single answer.
Context
It fits with what we’d already been noting about AI engines that disagree with each other when describing the same brand and about how machine-made content gets treated. The problem isn’t the accuracy of the facts; it’s choosing between defensible alternatives.
What’s next
- The paper has been on arXiv since September 22, without peer review.
- No timelines announced for releasing Taste-Bench as a public benchmark or for adding more task families.
Bottom line
The industry has spent two years debating whether models hallucinate. This measurement asks something else: if they get the facts right but choose badly, the text still comes out wrong.
Sources
Edited by Rodrigo Cornejo. How we select, verify and correct each note.


