A new study out of the UK is getting passed around education circles with a headline that sounds like every other AI-tutoring press release: an AI tutor matched — and on one measure, beat — human tutors. If you’ve read enough of these announcements, you know the pattern. A vendor runs a small study, cherry-picks the win, and quietly leaves out the part that actually explains the result.
This one is different, and it’s worth fifteen minutes of your attention, because the part everyone will bury in the headline is the part that matters most.
What the study actually found
Researchers from Eedi, a UK math learning platform, and Google DeepMind ran an exploratory randomized controlled trial across five UK secondary schools last summer. 165 students, ages 13-15, were split into groups: some worked with human tutors, some worked with LearnLM — Google DeepMind’s AI model built specifically for teaching, not general chat.
The results, on paper, look like a win for AI: students who worked with LearnLM solved new problems on later topics 66.2% of the time, compared to 60.7% for students who worked with human tutors alone. Across every other outcome the researchers measured, the AI-assisted group performed at least as well as the human-only group.
If that were the whole story, it would be one more entry in the pile of “AI replaces the tutor” claims — and it would fail Auxesis’s mission filter immediately. But it isn’t the whole story.
The detail the headline leaves out
Here’s what actually happened in that “AI tutoring” condition: a human expert tutor supervised every single message LearnLM drafted, in real time, before it reached a student. The tutor could edit it, replace it, or block it entirely. The AI wasn’t tutoring anyone on its own — it was drafting, and a human was editing.
The tutors approved 76.4% of LearnLM’s drafts with zero or minimal changes. That’s a real number worth sitting with: three out of four times, the AI got it right enough that a trained human, watching closely, didn’t need to intervene. That’s a legitimate, useful result. It means a well-designed AI model can shoulder real pedagogical weight.
But read that number the other way, too: nearly one in four times, a human needed to step in and change what the AI was about to say to a student. In an unsupervised deployment — the kind most homeschool families and classrooms actually get when they buy an AI tutoring product — nobody catches that one-in-four moment. The mistake just goes out.
This is Auxesis’s argument, run as a controlled experiment
We’ve said for a while now that the difference between AI that helps kids learn and AI that quietly hollows out their learning isn’t the model. It’s whether a human stays in the loop, watching what the AI actually does and stepping in when it drifts.
This study is that argument, tested under real classroom conditions with a real control group. It didn’t work because the AI replaced a tutor. It worked because a tutor was still there, reading every message, catching roughly one in four before a student ever saw it.
That’s not a footnote. That’s the mechanism.
It’s also, not coincidentally, close to how we’ve built COMPASS. Before a session, a Facilitation Brief gives the human facilitator context on where a student actually stands. During the session, a human is still the one teaching — reading the student, deciding what to say, when to push and when to pull back. After the session, a five-minute note captures what happened so the next session starts smarter. AI supports every stage of that loop. It never runs the loop alone.
What this means for you
If you’re a parent shopping for an AI tutoring tool, the Eedi study gives you a real, useful question to ask any vendor: Is a qualified human reviewing what this AI says to my child, in real time, before it says it? Or is the AI just talking directly to my kid, unsupervised?
Most consumer AI tutoring products on the market today are the second kind. This study didn’t test that kind. It tested the first kind — and even under close human supervision, with expert tutors catching nearly a quarter of the AI’s drafts, it still needed that oversight to be safe and effective.
If you’re an educator or an institution evaluating AI tools, the number to hold onto isn’t 66.2% vs. 60.7%. It’s 76.4%. That’s your actual design question: not “can AI tutor,” but “what’s your plan for the roughly one message in four that a human needs to catch?”
The Eedi researchers deserve real credit here — they published the number that most vendors would have left out. That’s the kind of study worth taking seriously, and the kind worth being honest about: it’s a genuine, well-designed result, and it’s also exploratory (N=165, one summer, five schools) — a strong first data point, not a settled verdict.
Sources: Eedi & Google DeepMind, “AI tutoring can safely and effectively support students” (arXiv): https://arxiv.org/abs/2512.23633. The 74: “AI Tutors, With a Little Human Help, Offer ‘Reliable’ Instruction, Study Finds”: https://www.the74million.org/article/ai-tutors-with-a-little-human-help-offer-reliable-instruction-study-finds/
This is a test to make sure everything is wired to capture information.