What happened
- A paper posted to arXiv on September 24 proposes ICLR, a method for discarding an agent’s accumulated reasoning without hurting its performance. It is signed by Mingxuan Wang, Fei Luo, Bo Wang and six other authors.
- The method needs no training and runs online. It scores each stretch of reasoning with an entropy measure computed by a frozen auxiliary model, and leaves actions, tool calls and observations untouched.
- Across 260 WorkBuddyBench tasks, average reward rises from 0.699 to 0.718 while input tokens (25.5%), output tokens (14.4%) and cache reads (33.3%) all fall.
- The authors describe an effect they call trajectory amplification: deleting reasoning at one point produces disproportionate changes in compute over the following interactions.
- Their conclusion: reasoning that has already been logged becomes replaceable once the derived information that matters has moved into code, files, tool outputs or responses from the environment.
Why it matters
- The cost of a long-horizon agent is dominated by the context it drags along, not by how hard the task is. Cutting input tokens by a quarter changes the arithmetic of any process that runs thousands of times a month.
- The condition the authors identify is the reusable part: if the conclusion has been written to a file or to a tool’s output, the reasoning that produced it can be thrown away. That is a design rule, not a compression trick, and it can be applied without adopting the method.
- For teams billed on inference consumption, the drop in cache reads is the most direct figure: 33.3% less on the line item that usually gets left out of early estimates.
The number
33.3% fewer cache reads, the largest reduction of the three measures.
Context
Context-compression strategies have so far summarized the full history, with the known problem that a summary can introduce errors the agent then treats as facts. This paper changes the criterion: instead of summarizing everything, it distinguishes which part of the history is working state and which part is a permanent record, and discards only the former.
What’s next
- The paper has been available on arXiv since September 24. No timelines have been announced for releasing the code or for tests in production environments.
Bottom line
An agent that remembers everything it thought pays to remember it at every subsequent step. This paper puts a number on something already suspected: much of that archive is useless once the result has been written down somewhere else.
Sources
- “When Can Agents Forget Their Reasoning?”, arXiv 2609.29875
- Our previous coverage: agents that report unfinished tasks as done, Docker’s isolated cloud environments and DeepSeek Flash’s time-of-day pricing.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


