Hallucinations

Models pile up the answer before doubting, and the correction comes too late


An analysis of ten models locates the components that push toward committing and the ones that hold back, and shows the latter always act afterward.

September 29, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

12.2 points of improvement in the accuracy of the decision to answer or abstain.

Context

Stanford had already measured hallucination rates from 22% to 94% and models that get worse when the user introduces the error. This paper goes one level down: instead of counting how often it fails, it describes where the decision that fails is made.

What’s next

Bottom line

The artificial intelligence scribes of the British health system confuse drugs, and it is patients who catch the errors. If the signal of doubt was inside and went unused, the problem was not that the model didn’t know. It was that no one asked it in time.

Sources


Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify each note.

Related notes

← All notes