What happened
- A paper announced on September 29 defines hallucination as unsupported commitment: the model answers with confidence despite showing internal signals that the question cannot be answered.
- Using causal gating, the authors identify a sparse, localized circuit, a subset of attention heads and middle layers, that governs the decision between answering and abstaining.
- The pattern repeats across ten models from 3 to 14 billion parameters, from five different families and on three benchmarks: the components that push toward commitment accumulate in early layers, and those that push toward abstention act later, as a correction that often fails to undo what has piled up.
- A lightweight policy trained on that circuit’s activations improves decision accuracy by 12.2 points, cuts false abstentions by a factor of 2.5, transfers to unseen benchmarks and scales to models of 27 to 35 billion parameters.
Why it matters
- The result reorders the problem. Hallucination stops being described as a lack of knowledge and starts being described as a problem of internal timing: the doubt exists, but it arrives late.
- Cutting false abstentions by 2.5 times matters as much as reducing invented answers. A system that refuses to answer too often is the one teams end up switching off after two months.
- For anyone deploying assistants in Chile over their own documentation, this supports a concrete design decision: the signal that the model doesn’t know exists and can be extracted, so reading it is worth more than asking the model to declare it in its answer.
The number
12.2 points of improvement in the accuracy of the decision to answer or abstain.
Context
Stanford had already measured hallucination rates from 22% to 94% and models that get worse when the user introduces the error. This paper goes one level down: instead of counting how often it fails, it describes where the decision that fails is made.
What’s next
- No announced date for releasing the causal-gating method or the trained policy.
- The authors report no results on closed models, where activations are not accessible. No timeline to extend it.
Bottom line
The artificial intelligence scribes of the British health system confuse drugs, and it is patients who catch the errors. If the signal of doubt was inside and went unused, the problem was not that the model didn’t know. It was that no one asked it in time.
Sources
- The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining, arXiv, announced September 29, 2026.
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.


