AI in health

Nine models leave out clinical information when reading full medical records


The benchmark was validated by 19 physicians and is renewed to prevent data leakage. The failures cluster in questions spanning several documents.

September 27, 2026 · Translated from the Spanish original

What happened

Why it matters

The number

19 physicians validated the automatically generated questions and answers.

Context

AI clinical transcription tools had already shown errors when deployed in public health services, made worse by the fact that the review falls to professionals with no time to do it. Here the object measured is different — retrieving existing information, not generating it — and the result points the same way: the error hides in what’s missing.

What’s next

Bottom line

A benchmark that renews itself so it can’t be memorized admits something uncomfortable about the state of evaluation: the useful life of a measurement is counted in months, and depends on when the set ends up in the next training corpus.

Sources

Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.

Related notes

← All notes