For two years, kids in 18 middle schools in Hamilton County, Tennessee, had access to Khanmigo — Khan Academy’s AI tutor — during their daily remedial math block. This wasn’t a loose, opt-in pilot. It was built into existing class time, and the AI was set up the “right” way: configured to coach, not to hand over answers. Ask a question back. Make the student do the thinking. It’s the exact design pattern Auxesis has argued for since day one.
Researchers Philip Oreopoulos (University of Toronto) and Nina Low (Charles River Associates) ran it as a real randomized trial and just published the results through the National Bureau of Economic Research. The honest number: assignment to Khanmigo raised math achievement by about 1.3 national percentile ranks per term — roughly 0.06 to 0.08 standard deviations over a school year. If you isolate students who actually engaged with it for a full year, the effect climbs to 0.14 standard deviations. That’s a real, positive number. It is also, in the authors’ own words, a gain that “resembles those from Khan Academy practice without AI assistance.” The AI barely moved the needle past what the plain practice platform already did.
So what happened? Not the questions. The questions were good — that was the whole point of the configuration. What happened was nobody showed up to answer them.
96% of students tried Khanmigo at least once. That’s a strong opt-in number — kids were curious, they clicked in. But the median student only messaged it on about one day in three that they were practicing math, and in just 17% of the sessions where they actually got something wrong — the exact moment a coaching question matters most. The researchers are direct about the cause: student engagement, not model capability, was the binding constraint.
This is worth sitting with, because it’s not a story about a bad AI tutor. It’s a story about a good one. Khanmigo did what it was told: it asked instead of told. And it still stalled at roughly the same outcome as a kid working through practice problems with no AI at all — because asking a good question into an empty room doesn’t teach anyone anything. Someone has to be there to notice the question got skipped, and to make sure it gets answered anyway.
That’s the piece every “coach, not answer” AI tutor is currently missing, and it’s the specific gap COMPASS — Auxesis’s own operating layer for facilitators — is built to close. The Read–Respond Loop is a three-part habit: before a session, the facilitator reads a short Facilitation Brief so they know exactly where a student is stuck; during the session, the teaching stays human — a real person watching whether the student engaged with the question or breezed past it; after, a five-minute note captures what actually happened, so the next session picks up exactly where this one left off. None of that requires a smarter model. It requires a person whose job is to notice when a student didn’t answer the question the software asked — and make them go back and answer it.
Picture the Tennessee classroom differently for a second. Same Khanmigo, same coaching configuration, same math block — but a facilitator trained to run the Read–Respond Loop is checking in on the kids who logged zero messages that week, not just the overall usage dashboard. That’s not a hypothetical improvement; it’s the missing half of exactly the experiment that just ran. The software held up its end. The open question this study leaves — who holds up the other end — is the one Auxesis’s Facilitator Certification exists to answer.
The AI tutoring industry is quietly converging on “ask, don’t tell” as the default design. That’s genuinely good news; it means the pedagogy Auxesis has taught from the start is being validated at scale, by people with no reason to agree with us on purpose. But this study is the clearest evidence yet that the design pattern alone isn’t the finish line. A tool that asks the right question and a person who makes sure it gets answered are two different jobs. Right now, most AI tutors are only staffed for the first one.
If you’re using an AI tutor with your kid or your students right now — do you actually know whether they’re engaging with it, or just opening it once and calling it done?

Leave a Reply