What happened
- On April 14, 2026, Stanford University’s Institute for Human-Centered Artificial Intelligence published the ninth edition of its annual report on the state of AI.
- In a new accuracy indicator, the hallucination rates of 26 leading models range from 22% to 94%.
- The report documents that GPT-4o’s accuracy fell from 98.2% to 64.4% on that indicator, and DeepSeek R1’s from more than 90% to 14.4%.
- The most specific finding is about framing: when a false claim is presented as something another person believes, the models handle it well. When it’s presented as something the user believes, performance collapses.
Why it matters
- Anyone writing with an assistant is in the worse of the two scenarios measured. When asking for help, you describe your own reading of the matter, and that is precisely the framing under which the model stops correcting. The tool goes along with the error instead of stopping it.
- The operational consequence is that reviewing the output isn’t enough. You also have to review the premise you put into the instruction, because the model had little incentive to challenge it.
- The 22% to 94% range undoes the idea that models are interchangeable. Between the best and worst in the sample there’s a gap of more than seventy points on the same indicator.
- Developers consistently publish capability results and very little about accuracy, bias or transparency. Whoever picks a vendor compares what’s published, and what’s published is the favorable half.
The number
362. AI incidents documented during 2025, compared with 233 in 2024.
Context
The report also records that AI-specific governance roles grew 17% in 2025 and that the share of companies with no responsible-use policy at all fell from 24% to 11%. The obstacles cited for implementing one are lack of knowledge (59%), budget (48%) and regulatory uncertainty (41%).
What’s next
- No timelines announced. The report comes out once a year; the tenth edition would be in 2027.
Bottom line
Recent research cited in the same document found that improving one dimension of responsible use consistently degrades another: training for safety costs accuracy. The trade-off isn’t resolved, and it’s already being sold as if it were.
