Why AI Math Tutors Keep Failing at Multi-Step Reasoning

Khan Academy built Khanmigo to be different. Instead of giving students answers, it was supposed to guide them — asking questions, nudging reasoning, mimicking what a great tutor would do in a one-on-one session.

It isn’t working. Khan Academy has since described Khanmigo’s actual classroom impact as a “non-event,” with usage sitting around 15% of eligible students. Khanmigo isn’t the only one struggling — Quizlet quietly retired its own AI tutor, Q-Chat, last year, without giving a detailed public reason, though the economics of running a generative tutor at a subscription price point are a plausible culprit.

Two of the best-funded AI tutoring bets in education have both quietly stumbled within the last year — one on adoption, one on the business model underneath it. That’s worth pausing on, because it isn’t a coincidence. Both tools were built on the same flawed assumption about what multi-step math reasoning actually requires.

The Assumption That Keeps Breaking

Both tools were designed around a chat interface: type a question, get a response, keep going. That works for recall — “what’s the formula for area?” — but multi-step reasoning isn’t a recall problem. It’s a sequencing problem. A student doesn’t fail at a two-step word problem because they don’t know a fact. They fail because they haven’t yet built the internal scaffolding to move from the concrete situation, to a picture of it, to the abstract equation that solves it.

A chat window can ask a leading question. It can’t watch a student’s confusion in real time and decide whether they need a manipulative in their hands, a bar model on paper, or one more nudge before the abstraction clicks. That’s not a prompt-engineering problem you patch with a better system prompt. It’s a structural mismatch between the medium (text chat) and the task (building number sense in stages).

What This Means If You’re Choosing a Tool Right Now

If you’re a homeschool parent or educator evaluating an AI math tool today, here’s the practical takeaway: don’t judge a tool by whether it “asks good questions.” Judge it by whether it follows a real sequence — concrete, then pictorial, then abstract — and whether it actually changes its approach when a student is stuck, rather than just rephrasing the same question in a friendlier tone.

This is exactly why Auxesis builds around Singapore Math’s Concrete-Pictorial-Abstract framework and Socratic, Zero-Reveal facilitation instead of a generic chat tutor. The goal was never to build a faster way to give answers. It’s to make sure a student can explain their reasoning, apply it to something new, and still have it a month later — the three conditions that separate real mastery from the appearance of it.

Khanmigo and Q-Chat aren’t failures of effort. They’re evidence that the first generation of “AI tutor as product” bets is running into the same wall from two directions — one on whether kids actually learn from it, one on whether anyone can afford to run it. Worth remembering the next time a tool promises to be the students’ patient, infinitely available tutor: patient and available isn’t the same as pedagogically sound.

A narrow question if you’re evaluating a tool right now: the last time your kid got stuck, did it change its approach — or just ask the same question again in a friendlier tone?

Leave a Reply

Your email address will not be published. Required fields are marked *