Agent costs

Deleting an agent's old reasoning cuts input tokens by 25.5%


The method needs no training and lifts average reward from 0.699 to 0.718 across 260 tasks, according to a paper published on September 24.

September 27, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

33.3% fewer cache reads, the largest reduction of the three measures.

Context

Context-compression strategies have so far summarized the full history, with the known problem that a summary can introduce errors the agent then treats as facts. This paper changes the criterion: instead of summarizing everything, it distinguishes which part of the history is working state and which part is a permanent record, and discards only the former.

What’s next

Bottom line

An agent that remembers everything it thought pays to remember it at every subsequent step. This paper puts a number on something already suspected: much of that archive is useless once the result has been written down somewhere else.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes