What happened
- On September 23, a group led by Curtis Northcutt published StudentBench on arXiv, a platform and evaluation suite for comparing human tutoring with language model tutoring.
- The study followed 2,383 participants preparing for the GRE and collected more than 175,000 interactions between students and AI systems.
- The central result, according to the paper itself: AI tutoring was statistically equivalent to expert human tutoring in learning gains, with p = 0.015. In five of the seven domains evaluated, the AI tutors outperformed the humans.
- The cost per percentage point gained was $0.0052 with AI and $4.81 with a human tutor. The authors also measured that shorter response times were associated with greater student participation.
Why it matters
- The finding isn’t that AI teaches better. It’s that it teaches just as well at a fraction of the cost, and that turns a pedagogical debate into a budget debate. Budget debates get settled faster.
- For education systems in the region, where private tutoring is an out-of-pocket expense that separates those who can afford it from those who can’t, the cost figure is the one that will move decisions. Preparation that costs cents per point redefines who gets access to a competitive score.
- There’s a limit the design itself imposes, and it’s worth stating: what was measured was preparation for a standardized test, with bounded content and known correct answers. Extrapolating that to school teaching is a claim the study doesn’t support.
The number
$0.0052 versus $4.81. That’s the cost of each percentage point of improvement, with AI or with a human tutor.
Context
The evidence on AI applied to people had been more ambiguous than enthusiastic. It has been documented that the effect of AI companions on well-being depends on usage and that much of the hallucination depends on how the user asks. StudentBench contributes what was missing from that debate: a measurement of outcomes, not perceptions.
What’s next
- The paper is a preprint, without peer review. The authors published the platform at studentbench.org.
- There’s no announcement of a replication in other languages or on tests other than the GRE.
- No date for journal publication.
Bottom line
The figure that will circulate is the 918-fold difference in cost. The one that decides anything is the other one: five of seven domains. An average that comes out even can hide subjects where the human tutor is still better, and the paper doesn’t say which two those were.
Sources
- StudentBench, arXiv 2609.28470, September 23, 2026.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


