What happened
- A paper announced on September 29 studies what happens in an extended autonomous research task after the person hands over the request and leaves.
- The method is counterfactual: a single user factor relevant to the task is changed, the rest of the context is held fixed, and the information requests, working drafts and final reports are compared.
- The central result is a separation between process and delivery. Agents research similar broad questions and split their requests differently, but it is the final recommendations that distinguish the user conditions most clearly.
- The reports also integrate user factors that were not visible at the same time during the search, and the recommendations keep their direction despite major rewrites between draft and report. The difference repeats across several models, runtime environments and evaluators.
Why it matters
- Auditing an agent’s search trail is not enough to know why it recommended what it recommended. The paper shows that the process looks alike across cases and the delivery does not.
- For someone who commissions a market report or a supplier review from an agent, this defines where to place human review: on the recommendation, not on the bibliography.
- It also defines a concrete commercial risk. If the recommendation changes according to a detail of the request that no one looked at again, responsibility for that bias falls on whoever wrote the brief.
The number
A single factor changed is enough to separate the final recommendations, with the rest of the context held fixed.
Context
OpenAI had already declared its goal of the automated research intern met, grading itself. Mathematicians, on the other hand, read a proof generated by those systems and learned nothing from it. The question of what remains auditable after autonomous work has been going around for a month.
What’s next
- The authors leave unresolved the ambiguous cases the method doesn’t classify. No date for a version that addresses them.
- No announced timeline for releasing the evaluation framework.
Bottom line
A recent method showed that deleting an agent’s old reasoning cuts input tokens by 25.5% without losing performance. If what gets deleted wasn’t what determined the recommendation, the question stops being how much context to keep. It becomes which.
Sources
- What Happens During Autonomous Deep Research After the User Steps Away?, arXiv, announced September 29, 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.


