What happened
- A paper posted to arXiv on September 24 presents BRIE, a benchmark for measuring how language models retrieve information when they read electronic health records. It is signed by 26 authors, with Jordan L. Cahoon first and Emily Alsentzer last.
- The framework automatically generates question-and-answer pairs from longitudinal clinical notes. Nineteen physicians validated the result.
- Nine language models were evaluated with five different inference strategies.
- The central finding: state-of-the-art systems frequently leave out clinically important information, especially in questions that require synthesizing several documents and several encounters.
- Unlike static benchmarks, BRIE accepts several valid answers for the same question — because clinical reasoning varies among physicians — and renews its content continuously so it doesn’t end up inside training data.
Why it matters
- The failure mode is the worst possible one for a health system: it doesn’t answer wrong, it answers short. An omission doesn’t trigger any alarm in the workflow, and there’s no way to tell it apart from a record that simply didn’t contain the data.
- The concentration of errors in multi-document questions points straight at the commercial promise being sold to healthcare providers in the region: summarizing a patient’s full history. That is exactly the task where the evaluated models fail most.
- For anyone buying these tools, the design of the benchmark matters as much as the result. A set that renews itself and accepts multiple correct answers is the only one that still measures something six months later; accuracy figures on static tests age without warning.
- In Chile, medical records are sensitive data under Law 21,719. This note describes technical evidence and does not constitute legal advice.
The number
19 physicians validated the automatically generated questions and answers.
Context
AI clinical transcription tools had already shown errors when deployed in public health services, made worse by the fact that the review falls to professionals with no time to do it. Here the object measured is different — retrieving existing information, not generating it — and the result points the same way: the error hides in what’s missing.
What’s next
- The paper has been available on arXiv since September 24. No timelines announced for opening the benchmark to the public or for a results table by model.
Bottom line
A benchmark that renews itself so it can’t be memorized admits something uncomfortable about the state of evaluation: the useful life of a measurement is counted in months, and depends on when the set ends up in the next training corpus.
Sources
- “A Living Benchmark for Information Retrieval from Electronic Health Records”, arXiv 2609.30205
- Our previous coverage: the errors of AI scribes in the British health system, the hallucinations users don’t detect and the consequences of postponing the data protection law.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


