Schools Are Buying AI By the Billions. Nobody Can Prove Any of It Works.
Last year, U.S. schools spent about $2.5 billion on AI tools. By 2033, that number is expected to top $15 billion. Somewhere in a district office near you, someone is…
Last year, U.S. schools spent about $2.5 billion on AI tools. By 2033, that number is expected to top $15 billion. Somewhere in a district office near you, someone is deciding which of those tools gets a contract — and according to a new Stateline investigation, they’re often doing it with less information than the company selling them the product.
“There’s always an asymmetry between what the providers know and what the districts know,” one official told Stateline. Ed-tech purchasing has always had this problem. What’s different with AI is speed: the tools are changing faster than any district’s ability to evaluate them, and unlike a textbook, you often can’t just flip through an AI tool and see what it actually does.
This isn’t a one-off complaint. Instructure’s 2026 Evidence Report — a survey of classroom technology broadly, not just AI — found that most consumer ed-tech in schools today lacks any verified proof that it actually helps students learn. Tools get adopted because they’re popular, because a rep gave a good demo, or because “everyone else is using it.” Whether they work is a separate question nobody’s actually answering.
A case study in what “no way to verify” looks like
You don’t have to look far for a concrete example. Alpha School — the AI-driven “two-hour learning” model now expanding toward 50 campuses nationwide — has built its entire pitch around a headline claim: students learn twice as fast.
That number comes entirely from Alpha’s own analysis of its students’ NWEA MAP Growth scores. Alpha submits the data. NWEA records it. Alpha publicizes it. NWEA itself hasn’t audited it, no independent researcher has been given access to check it, and there’s no control group to compare against. According to Tech Times, Alpha’s own team promised to hand over raw data files for an independent audit — as of this writing, three months after that promise, the files still haven’t arrived.
To be clear: this isn’t an accusation that the number is false. It’s the opposite problem — nobody outside Alpha is in a position to say whether it’s true. And that’s exactly the gap Stateline and Instructure are describing at the scale of an entire industry, playing out in one specific, checkable claim.
What “verified” would actually require
If a district can’t tell whether an AI tool works, the honest answer isn’t “trust the vendor” or “wait for a study that may never come.” It’s building the evaluation into how the tool gets used in the first place — so the evidence exists whether or not anyone later asks for it.
That’s the idea behind COMPASS, Auxesis’s AI operating system for facilitated learning. Before a session, a Facilitation Brief tells the human teacher or tutor what to focus on. During the session, a person — not the AI — does the actual teaching. After, a five-minute note captures what happened, in plain language, tied to a real student and a real concept. None of that data lives in a black box. It’s not a marketing claim aggregated at the end of the year; it’s an ongoing, readable record of what was taught, to whom, and whether it landed — built for a parent, a teacher, or a district to actually look at.
That’s a smaller promise than “twice the learning.” It’s also one you can check.
The real question for anyone buying AI right now
Before any district signs a contract, before any parent pays for a subscription, the Stateline and Instructure findings point to the same practical question: not “does this claim sound impressive,” but “can I see the evidence myself, or am I taking the vendor’s word for it?”
Auxesis exists because we think that question deserves a real answer — not a promise that’s still three months overdue.
If your district or your kid’s school is using an AI tool right now — do you actually know what evidence backs it up, or are you taking it on faith? We’re curious what people have found when they’ve asked.
Why AI Math Tutors Keep Failing at Multi-Step Reasoning
Khan Academy built Khanmigo to be different. Instead of giving students answers, it was supposed to guide them — asking questions, nudging reasoning, mimicking what a great tutor would do in a one-on-one session.
It isn’t working. Khan Academy has since described Khanmigo’s actual classroom impact as a “non-event,” with usage sitting around 15% of eligible students. Khanmigo isn’t the only one struggling — Quizlet quietly retired its own AI tutor, Q-Chat, last year, without giving a detailed public reason, though the economics of running a generative tutor at a subscription price point are a plausible culprit.
Two of the best-funded AI tutoring bets in education have both quietly stumbled within the last year — one on adoption, one on the business model underneath it. That’s worth pausing on, because it isn’t a coincidence. Both tools were built on the same flawed assumption about what multi-step math reasoning actually requires.
The Assumption That Keeps Breaking
Both tools were designed around a chat interface: type a question, get a response, keep going. That works for recall — “what’s the formula for area?” — but multi-step reasoning isn’t a recall problem. It’s a sequencing problem. A student doesn’t fail at a two-step word problem because they don’t know a fact. They fail because they haven’t yet built the internal scaffolding to move from the concrete situation, to a picture of it, to the abstract equation that solves it.
A chat window can ask a leading question. It can’t watch a student’s confusion in real time and decide whether they need a manipulative in their hands, a bar model on paper, or one more nudge before the abstraction clicks. That’s not a prompt-engineering problem you patch with a better system prompt. It’s a structural mismatch between the medium (text chat) and the task (building number sense in stages).
What This Means If You’re Choosing a Tool Right Now
If you’re a homeschool parent or educator evaluating an AI math tool today, here’s the practical takeaway: don’t judge a tool by whether it “asks good questions.” Judge it by whether it follows a real sequence — concrete, then pictorial, then abstract — and whether it actually changes its approach when a student is stuck, rather than just rephrasing the same question in a friendlier tone.
This is exactly why Auxesis builds around Singapore Math’s Concrete-Pictorial-Abstract framework and Socratic, Zero-Reveal facilitation instead of a generic chat tutor. The goal was never to build a faster way to give answers. It’s to make sure a student can explain their reasoning, apply it to something new, and still have it a month later — the three conditions that separate real mastery from the appearance of it.
Khanmigo and Q-Chat aren’t failures of effort. They’re evidence that the first generation of “AI tutor as product” bets is running into the same wall from two directions — one on whether kids actually learn from it, one on whether anyone can afford to run it. Worth remembering the next time a tool promises to be the students’ patient, infinitely available tutor: patient and available isn’t the same as pedagogically sound.
A narrow question if you’re evaluating a tool right now: the last time your kid got stuck, did it change its approach — or just ask the same question again in a friendlier tone?
Picture a Florida school district in October. A teacher has been using a reading-support tool for three years – something that adapts to how a struggling reader sounds out words, nothing flashy, just quietly effective. A parent has been using a Socratic AI tutor with her homeschool co-op since spring, the kind that asks questions instead of giving answers. Neither one has ever thought of what they’re using as a chatbot. They think of it as a tool that helps a kid learn.
Then, on August 5, 2026, the Florida Department of Education sat down for a rule-development workshop. The goal was simple and, honestly, overdue: protect kids from AI chatbots that simulate friendship, that keep a lonely teenager talking longer than is healthy, that were never built with a classroom in mind. Nobody in that room was trying to make life harder for the reading tool or the Socratic tutor.
But here’s what happened. The draft language – an amendment to the state’s Internet Safety Policy, Rule 6A-1.0957 – defined “artificial intelligence” broadly enough to catch everything with a chat interface in the same net. The vetted reading tool. The homeschool tutor. The companion chatbot the rule was actually written for. All three, same definition, same rules: a parent has to opt in before their kid can use it, and the school has to offer a non-AI alternative assignment instead.
The Software and Information Industry Association (SIIA) – the trade group representing the very companies that build these tools – read the draft and said, in effect: you’ve aimed at the wrong target. Their public comment warned that pairing this broad definition with the opt-in and alternative-assignment requirement would bury schools in paperwork, and would hit hardest exactly the kids who benefit most from adaptive tools – struggling readers, English language learners, students with disabilities. The people the rule was supposed to protect are the people most likely to get cut off by it.
This is still a draft. Nothing is final. But there’s a real date attached to it: under the current version, Florida’s districts and charter schools have until January 1, 2027 to adopt whatever policy comes out the other side of this process. That’s not far off. (Worth knowing, separately, that Florida also has SB 482, the “Artificial Intelligence Bill of Rights,” working through the legislature – a different bill, a different track, easy to confuse with this one but not the same fight.)
Here’s the part that matters beyond Florida. This is the shape 2026’s AI-in-education rulemaking keeps taking, state after state: good instincts, broad language, and a definition that can’t yet tell the difference between a tool that replaces a child’s thinking and one that builds it. If you’re running a microschool, a co-op, or a tutoring program with any AI-assisted piece – a chat interface, an adaptive practice tool, anything that talks back – “we’re not a companion chatbot” won’t hold up once a rule like this is final. You need to be able to show the difference, not just claim it.
That’s the same problem COMPASS was built to solve. A facilitation brief. A session log. A record of what the AI did and what the human did, in plain language, that a regulator can actually read. Programs that can produce that evidence will clear whatever definition Florida – or the next state – eventually settles on. Programs that can’t will be scrambling in December, trying to prove after the fact something they should have been documenting all along.
Florida hasn’t finished writing this rule yet. That’s the opportunity. The five months between now and January 2027 are exactly the runway to make sure your own program can prove, not just assert, that what you’re doing is augmenting a kid’s thinking – not replacing it.
Sources: SIIA public comment on Florida DOE Rule 6A-1.0957 rulemaking; Florida DOE rule-development workshop record, Aug 5 2026; independent legislative tracking of SB 482 / HB 659.
On August 12, a new kind of private school opens its doors in Oklahoma. Alpha School – a billionaire-backed chain that started in Austin in 2014 – is launching a K-10 campus in Edmond and a K-8 campus in Tulsa. Tuition is $40,000 a year, nearly double the most expensive private school in either metro area. There are no teachers.
That’s not a figure of speech. Alpha’s model replaces teachers with two hours a day of AI-driven, self-paced instruction – 30 minutes each for math, reading, science, and social studies – followed by afternoon workshops in life skills and entrepreneurship. The adults on staff are called “guides,” and by design, they don’t teach. According to Alpha’s own job postings, guides monitor screens and run afternoon workshops, but they never touch curriculum, and they don’t need a teaching certificate to get hired. One guide profiled in coverage of the school was a former sports psychologist.
We think this is worth paying attention to – not because Alpha is unusual (it isn’t; it’s opening 28 new campuses this fall, more than tripling its footprint), but because it is the clearest real-world test yet of the question at the center of everything we do: what happens when you take the adult who assesses, questions, and responds to a struggling kid, and replace them entirely with software?
What Alpha actually promises
To be fair to the model: Alpha isn’t hiding what it does. Families know exactly what they’re signing up for, and some are genuinely enthusiastic. One incoming Tulsa parent told Oklahoma Watch she researched every criticism of Alpha she could find before enrolling her son and “was not deterred.” Alpha reports strong outcomes – an average SAT score of 1545 out of 1600 among its seniors, and AP scores averaging 4.2-4.4. If those numbers hold up under independent scrutiny, they’re real.
But strong test scores and an unsupervised AI tutor are not the same evidence. A high score can mean a student learned the material – or that a system optimized for the test got them there without much staying power. That’s the distinction our whole approach is built around: augmenting the thinking, not replacing the person who checks whether the thinking actually happened.
The part that should give any parent pause
Two details from the reporting stand out.
First, the enrollment numbers. As of July 20 – three weeks before opening day – only 26 students had enrolled at the Edmond campus and 32 at Tulsa. That’s a small, self-selecting group of early adopters, not yet a broad verdict from Oklahoma families.
Second, the privacy design. Alpha’s AI platform monitors students’ screens continuously and takes photos when it detects “unusual screen patterns” or a lack of progress. In one case reported by Wired, the system captured a photo of a student doing schoolwork in bed, in pajamas. Separately, 404 Media and Wired both reported that Alpha’s underlying AI system may have been built in part by scraping content from other online learning platforms without permission – a claim serious enough that a line of broken contracts and licenses has followed the company. To be clear about where this stands: Alpha’s chief communications officer disputes the scraping claim directly, calling the reporting “cherry-picked” and sourced to former employees, and says the company may pursue legal action against the outlets that reported it. The screen-monitoring and photo-capture practice itself is not seriously disputed – it’s simply how the system is designed to work.
State lawmakers have noticed the staffing model specifically. Oklahoma Rep. Cyndi Munson put it directly: “These types of schools and models that are being implemented are taking away the seriousness of what it takes to be a classroom teacher. It is very disturbing and should worry all of us.” Erika Wright, founder of the Oklahoma Rural Schools Coalition, went further: “This is a travesty for those kids.” Alpha’s students, she argues, won’t get the educational and emotional support that comes from a certified adult who actually teaches them.
There’s a regulatory wrinkle here too. Oklahoma’s new AI-in-schools law – the Responsible Technology in Schools Act – requires public school districts to keep a human educator in final decision-making authority over any AI-assisted instruction. Alpha, as a private school, isn’t covered by that law at all. And in a detail that says a lot about how these programs get vetted, Alpha’s Edmond campus was briefly listed as approved for Oklahoma’s Parental Choice Tax Credit – a state voucher program – before the Oklahoma Tax Commission quietly removed it from that list after Oklahoma Watch started asking questions.
The mission-filter test
Run Alpha through the two questions we ask about anything AI touches in education:
Does this augment intelligence, or replace it? By design, Alpha replaces the person whose job is to notice when a kid is confused, reasoning incorrectly, or checked out – and substitutes a system that measures screen behavior instead.
Does this deepen learning, or shortcut it? That one we genuinely don’t know yet. Alpha’s test scores are real, if they hold up. But a test score tells you what a student can produce under exam conditions – not whether the understanding underneath it will still be there in six months, or whether a kid ever got the kind of “wait, why did you get that answer?” conversation that actually builds reasoning. Two hours of adaptive drilling a day, with no adult required to notice whether it’s landing, is a design that we’d expect to compress the appearance of mastery – the exact failure mode we spend the most time helping families and educators avoid.
We’re not the ones who get to answer that question about Alpha with certainty. But it’s the right question to be asking before $40,000 a year and a state voucher program get anywhere near a decision this consequential – and it’s the same question worth asking about any tool, anywhere, that promises to make learning faster by taking the human out of the loop.
Sources: KGOU/Oklahoma Watch, “A $40,000-per-year AI school with no teachers is opening in Oklahoma this August,” by Maya Henry, July 21, 2026. https://www.kgou.org/education/2026-07-21/a-40-000-per-year-ai-school-with-no-teachers-is-opening-in-oklahoma-this-august
Most debates about AI tutoring get stuck asking the wrong question: does AI help kids learn, or does it hurt? A new randomized controlled trial out of Texas A&M throws out that binary and tests something much closer to what we actually do at Auxesis — not “AI or no AI,” but what kind of AI, used how.
The setup
Researchers ran 90 tenth-grade science students through one of three conditions: no AI at all, a structured inquiry method called Argument-Driven Inquiry (ADI) on its own, or that same structured inquiry paired with a Socratic-dialogue AI tool (ChatGPT’s Study Mode, configured to ask guiding questions rather than hand over answers). Before and after, students were tested on scientific argumentation, critical thinking, self-efficacy, cognitive engagement, and metacognitive self-regulation.
What they found
The group using structured inquiry plus Socratic-dialogue AI came out ahead of both other groups — significantly greater gains in scientific argumentation, critical thinking, self-efficacy, and engagement. Not “AI helped a little.” A real, measurable edge, in a live science classroom, over both the no-AI control and the same instructional method without AI.
That’s worth sitting with, because it’s close to a direct test of the thing Auxesis is built around: AI that asks instead of tells, layered onto real instruction rather than standing in for it. Our Mastery Gates and Read-Respond Loop exist because we believe that “ask, don’t tell” mechanism is where the learning actually happens. This study didn’t set out to test Auxesis’s method specifically — but the architecture it tested is close enough to matter.
Two things we’re not going to skip past
First: this is a preprint, posted to Research Square in December 2025. It has not yet gone through peer review. That doesn’t make the finding wrong, but it hasn’t cleared the bar “published, peer-reviewed research” implies — better to say that plainly than let the framing imply more certainty than the study has earned.
Second: not every measure moved. Effects on metacognitive self-regulation — students’ ability to monitor and adjust their own thinking — were not statistically significant. The Socratic-AI group didn’t show a meaningful edge there. That’s a real result worth naming, not quietly leaving out. One guided-AI intervention, one semester, doesn’t automatically build every kind of self-directed thinking skill.
Why this matters for how you teach or facilitate
If you’re an educator experimenting with AI in your classroom — or building a facilitation practice through the Educator Track — this study is a useful anchor for two reasons. It’s rare to find a classroom RCT that isolates “Socratic vs. everything else” this cleanly, rather than lumping all AI use into one bucket. And its honest limitation (the metacognition null result) is a reminder that no tool, guided or not, does all the work of building a thinking student. That’s a facilitation job, not a software job — exactly the distinction COMPASS and the Facilitator Certification are built to keep front and center.
The bottom line
A well-designed AI tool that asks rather than answers, used inside real instruction, produced measurably better critical thinking and argumentation outcomes than either no AI or the same instruction without AI — in a live classroom, with 90 real students. That’s a strong data point for “augment, don’t replace.” It’s not the final word — one preprint, one grade level, one subject, one semester — and it didn’t move every measure. But it’s the clearest experimental support we’ve seen yet for the specific mechanism Auxesis is built around.
When you picture “AI in the classroom,” is your gut reaction hopeful or worried? Reply H or W — we’re curious where people land.
A new study out of the UK is getting passed around education circles with a headline that sounds like every other AI-tutoring press release: an AI tutor matched — and on one measure, beat — human tutors. If you’ve read enough of these announcements, you know the pattern. A vendor runs a small study, cherry-picks the win, and quietly leaves out the part that actually explains the result.
This one is different, and it’s worth fifteen minutes of your attention, because the part everyone will bury in the headline is the part that matters most.
What the study actually found
Researchers from Eedi, a UK math learning platform, and Google DeepMind ran an exploratory randomized controlled trial across five UK secondary schools last summer. 165 students, ages 13-15, were split into groups: some worked with human tutors, some worked with LearnLM — Google DeepMind’s AI model built specifically for teaching, not general chat.
The results, on paper, look like a win for AI: students who worked with LearnLM solved new problems on later topics 66.2% of the time, compared to 60.7% for students who worked with human tutors alone. Across every other outcome the researchers measured, the AI-assisted group performed at least as well as the human-only group.
If that were the whole story, it would be one more entry in the pile of “AI replaces the tutor” claims — and it would fail Auxesis’s mission filter immediately. But it isn’t the whole story.
The detail the headline leaves out
Here’s what actually happened in that “AI tutoring” condition: a human expert tutor supervised every single message LearnLM drafted, in real time, before it reached a student. The tutor could edit it, replace it, or block it entirely. The AI wasn’t tutoring anyone on its own — it was drafting, and a human was editing.
The tutors approved 76.4% of LearnLM’s drafts with zero or minimal changes. That’s a real number worth sitting with: three out of four times, the AI got it right enough that a trained human, watching closely, didn’t need to intervene. That’s a legitimate, useful result. It means a well-designed AI model can shoulder real pedagogical weight.
But read that number the other way, too: nearly one in four times, a human needed to step in and change what the AI was about to say to a student. In an unsupervised deployment — the kind most homeschool families and classrooms actually get when they buy an AI tutoring product — nobody catches that one-in-four moment. The mistake just goes out.
This is Auxesis’s argument, run as a controlled experiment
We’ve said for a while now that the difference between AI that helps kids learn and AI that quietly hollows out their learning isn’t the model. It’s whether a human stays in the loop, watching what the AI actually does and stepping in when it drifts.
This study is that argument, tested under real classroom conditions with a real control group. It didn’t work because the AI replaced a tutor. It worked because a tutor was still there, reading every message, catching roughly one in four before a student ever saw it.
That’s not a footnote. That’s the mechanism.
It’s also, not coincidentally, close to how we’ve built COMPASS. Before a session, a Facilitation Brief gives the human facilitator context on where a student actually stands. During the session, a human is still the one teaching — reading the student, deciding what to say, when to push and when to pull back. After the session, a five-minute note captures what happened so the next session starts smarter. AI supports every stage of that loop. It never runs the loop alone.
What this means for you
If you’re a parent shopping for an AI tutoring tool, the Eedi study gives you a real, useful question to ask any vendor: Is a qualified human reviewing what this AI says to my child, in real time, before it says it? Or is the AI just talking directly to my kid, unsupervised?
Most consumer AI tutoring products on the market today are the second kind. This study didn’t test that kind. It tested the first kind — and even under close human supervision, with expert tutors catching nearly a quarter of the AI’s drafts, it still needed that oversight to be safe and effective.
If you’re an educator or an institution evaluating AI tools, the number to hold onto isn’t 66.2% vs. 60.7%. It’s 76.4%. That’s your actual design question: not “can AI tutor,” but “what’s your plan for the roughly one message in four that a human needs to catch?”
The Eedi researchers deserve real credit here — they published the number that most vendors would have left out. That’s the kind of study worth taking seriously, and the kind worth being honest about: it’s a genuine, well-designed result, and it’s also exploratory (N=165, one summer, five schools) — a strong first data point, not a settled verdict.
Sources: Eedi & Google DeepMind, “AI tutoring can safely and effectively support students” (arXiv): https://arxiv.org/abs/2512.23633. The 74: “AI Tutors, With a Little Human Help, Offer ‘Reliable’ Instruction, Study Finds”: https://www.the74million.org/article/ai-tutors-with-a-little-human-help-offer-reliable-instruction-study-finds/
There’s no national policy for how AI should be used in schools. Not one. States are passing their own patchwork of rules, districts are writing their own, and most families are left guessing.
So a group of teenagers did what Congress hasn’t: they wrote the bill themselves.
On a late-July weekend, 98 teens from all 50 states gathered at the Edward M. Kennedy Institute for the United States Senate — inside a full-scale replica of the actual Senate chamber — and spent the weekend debating, amending, and voting on legislation to govern AI in K-12 schools. They called it S 2026, the Students First Act. It passed 83 to 15.
Here’s the part worth sitting with: the student-focused section of that bill has 15 provisions, and several of them circle back to the same worry, again and again, that AI is doing their thinking for them.
That’s not a rule an ed-tech company would write. It’s not a rule most school boards have written either. It’s a rule from the people actually sitting in the classroom, who can apparently tell the difference between AI that helps them think and AI that thinks for them, even when the adults writing policy around them mostly can’t.
Why This Matters More Than Another AI-in-Schools Headline
At Auxesis, we’ve spent a long time naming this exact problem: the illusion of mastery. It’s what happens when a student looks like they’ve learned something, the essay reads fine, the answer is correct, but the actual thinking never happened. Good output, no understanding underneath it.
Usually we’re the ones making that case, citing a study or a researcher. This time, 98 teenagers made it themselves, from the inside, without anyone telling them to. They didn’t call it the illusion of mastery. They called it AI doing their thinking for them. Same problem, plainer words.
That cuts against the two lazy stories adults tend to tell about kids and AI: either kids will use it responsibly with no guardrails needed, or kids can’t be trusted with it at all. The teens at the Kennedy Institute didn’t take either position. They drew a real line, AI for editing and brainstorming, yes, AI for the writing itself, no, the same kind of line Auxesis draws with the CPA framework (Concrete, Pictorial, Abstract) in math: AI can support you at any stage, but it can’t do the stage for you. Skip the struggle, and the mastery you end up with is fake, no matter how correct the final answer looks.
One of the provisions the teens landed on: no AI for writing. Editing, brainstorming, and studying with AI would be allowed starting in eighth grade — but the writing itself has to stay theirs.
What This Means for You
If you’re a homeschool parent, this is permission to trust your instinct that letting AI just explain it and letting AI help you think it through are not the same thing, a group of teenagers just spent a weekend proving they know that too.
If you’re an educator, this is a template. The Students First Act draws a line by task, not by blanket ban: editing and brainstorming are fine, the core cognitive work isn’t. That’s a more useful starting point for your own classroom AI policy than most of what’s floated around this year, and it’s the same logic behind Auxesis’s Facilitator Certification, which is built around exactly this judgment call: knowing when AI belongs in the room and when it doesn’t.
There’s still no national policy. But for one weekend, in a replica Senate chamber in Boston, 98 teenagers showed they’d already figured out what the adults are still arguing about.
Source: NPR, July 30, 2026 (npr.org/2026/07/30/nx-s1-5853571/students-set-ai-policy)
Team the 98 teens’ rule, AI for editing and brainstorming, never for the writing itself, or team stricter: no AI at all until high school? Curious which side your house or classroom lands on.
Want a framework for making this call yourself? The Parent Track and Educator Track both walk through exactly where AI belongs, and where it doesn’t.
Two weeks ago, a new AI tutoring company called Bloomy launched out of Y Combinator with a number attached: students in an early pilot grew 1.8 times faster than expected on the NWEA MAP assessment, a widely used test of academic growth. The founder has been careful — he’s called it an observational pilot, not a randomized study, and the sample was about 150 middle schoolers at one Massachusetts charter school.
That caveat matters, and other outlets have already covered it well: no control group, no randomization, one school. If you want the stats critique, it’s out there.
We want to talk about something the coverage has missed — because it’s sitting in Bloomy’s own product design, not in a footnote.
The part of Bloomy that looks a lot like us
Bloomy’s platform moves students through three stages for every skill: Base Camp (worked examples), Climb (guided practice with an AI tutor), and Summit — an unaided, ten-question test with the tutor deliberately absent. Students don’t advance until they clear roughly 90% on Summit.
If that structure sounds familiar, it should. It’s the same logic behind our Mastery Gates: Explain, Apply, Sustain. A student doesn’t move forward because they got through the material. They move forward because they can perform without support, cold, when nobody’s coaching them through it.
Here’s the question that matters: why build that gate at all?
Because a growth number doesn’t tell you what you think it tells you
NWEA MAP growth is a well-respected, widely used measure — but it measures performance on a norm-referenced test. It tells you a student answered more items correctly, or harder items correctly, than before. It does not, by itself, tell you why — whether a student built durable understanding, or got faster at pattern-matching the kind of problems that show up on that kind of test
That gap between “the score went up” and “the understanding is real” is what we call the illusion of mastery: the condition where a student looks like they’ve learned something, but the underlying thinking never actually happened. Good grades without reasoning. A rising RIT score without a student who can explain what they did.
The interesting thing is that Bloomy’s own architecture seems to already know this. If a rising test score were sufficient proof of learning, you wouldn’t need a Summit gate. You’d just assign more problems and let the score climb. The fact that Bloomy built an unaided, high-bar checkpoint before letting a student move on suggests someone on that team understands the difference between “the number went up” and “the student can actually do it alone” — even while the marketing leans on the number.
What this means for you
If you’re evaluating any AI tutoring tool — Bloomy or otherwise — for your family or your classroom, the growth-percentage headline is the least useful thing to ask about. The better questions:
Does the tool require a student to perform without help before it lets them move forward?
Can you see why a student got something wrong — a specific misconception — not just that they got it wrong?
Does “mastery” mean “answered correctly,” or does it mean “can explain the reasoning, apply it somewhere new, and still have it three weeks later”?
That’s the standard we hold our own Mastery Gates to, and it’s the standard worth holding any tool to — including ours.
If you’ve spent any time in parenting or education circles this year, you’ve seen the headlines: AI is making kids worse at thinking. Teachers can’t get students to reason anymore. There’s a new term for it – “the great unwiring” – and the evidence being passed around usually points to one study: nearly 1,000 high school math students in Turkey, split into groups, given access to ChatGPT-4 while they studied. The unrestricted group did 48% better on practice problems. Then, when the AI was taken away for the final exam, that same group scored 17% worse than students who’d never used AI at all.
That statistic is real. It comes from a peer-reviewed study published in PNAS – “Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics” (Bastani, Bastani, Sungu, Ge, Kabakci, and Mariman) – and it’s the study most people are pointing to when they cite this fear. It’s part of a broader wave of alarm: Fortune has covered a separate Brookings report warning of a “great unwiring” of students’ brains, and Axios has covered fresh polling showing most teachers think AI is hurting kids’ critical thinking. Every version of the story – whichever one you read – stops at the scary number.
Almost none of them mention the group that changes the whole story.
The Detail Getting Cut From Every Retelling
The researchers didn’t test one AI condition. They tested three: a control group with only textbooks and notes, a group with unrestricted ChatGPT-4 access (“GPT Base”), and a third group using a version of GPT-4 built with teacher-designed guardrails that gave hints and guiding questions instead of direct answers (“GPT Tutor”).
The GPT Base group is the one making headlines – 48% better on practice, 17% worse on the final exam. Used as a crutch, then gone, and the students had nothing underneath them.
The GPT Tutor group told a different story entirely. On practice problems, they didn’t just match the Base group – they outperformed it by 127%. And when the AI was removed for the final exam? They scored about the same as the control group. No penalty. No collapse. The guardrails held.
Synthesis Tutor calls itself, right there in the page title, “the world’s first superhuman math tutor.” That’s not marketing shorthand somebody exaggerated in an ad — it’s the actual tagline on their website today.
It’s a big claim. So when a claim is that big, the useful question isn’t “does that sound impressive?” It’s “what’s the evidence, exactly, and does it hold up?”
Synthesis actually makes this easy. They wrote a blog post called “Does the Synthesis Tutor get results?” and it names their source directly: a DARPA-funded program called the Digital Tutor. Follow that thread, and here’s what you find.
What the DARPA study actually was
The Digital Tutor was real, and by the numbers reported, it was genuinely impressive — for what it was designed to do. DARPA wanted to see how close a piece of software could get new recruits to the expertise of a seasoned professional, fast. So they built a tutoring system and tested it on U.S. Navy sailors training to become Information Systems Technicians — the people who keep a ship’s IT systems running.
Over 16 weeks, those sailors went through the Digital Tutor program. At the end, they were tested against two other groups: Fleet technicians with an average of ten years on the job, and sailors trained the traditional classroom way. The Digital Tutor group outperformed both, by a wide margin, on troubleshooting real IT systems.
That’s a strong result. It’s also, if you look at the actual DTIC records, from assessments run in 2010 — sixteen years ago now — on adult Navy technicians learning enterprise IT systems. Not elementary students. Not math. Not children at all.
Where the gap is
None of this means Synthesis Tutor is bad, or that the company is lying. They’re transparent about their source — they link straight to the DARPA documentation and a public summary of it. That’s more citation than most ed-tech marketing bothers with.
But there’s a real distance between “a 2010 program that helped Navy sailors master IT troubleshooting faster than classroom training” and “the world’s first superhuman math tutor” for a 6-year-old learning to subtract. The first is a specific, well-documented result in a narrow, adult, technical-skills domain. The second is a sweeping claim about a completely different subject, age group, and learning context — resting on that same study as its evidence.
That gap is the actual lesson here, and it’s a useful one to practice noticing — for you and for your kids.
A framework for checking claims like this
This is exactly the kind of moment where the third piece of metacognition — Evaluate — earns its keep. Monitor asks “do I understand this?” Regulate asks “what should I do differently?” Evaluate asks the question most of us skip: “was that reasoning actually sound?”
Applied to a marketing claim, Evaluate looks like three quick questions:
What’s the actual claim? (“Superhuman” — a specific comparative claim, not just enthusiasm.)
What’s the cited evidence? (A study — good, that’s better than nothing.)
Does the evidence match the population and the product? (Adult Navy IT trainees, 2010, technical troubleshooting — vs. a math app for a 7-year-old, in 2026.)
That third question is where most impressive-sounding claims quietly fall apart. It’s not about catching companies in lies — it’s about training the habit of checking whether the receipt matches the bill. That’s a skill worth modeling out loud with your kids the next time an app, a toy, or a tutor promises something “revolutionary.”
The takeaway
Citing a source is good. Citing the right source is what actually matters. Before any tool — ours included — earns a claim like “this works,” it should be able to show you evidence that was actually about the kids, the subject, and the context you’re deciding for. If it can’t, that’s worth knowing before you buy in.
Auxesis doesn’t sell superhuman anything. We teach the slower thing that actually works: struggle, checked understanding, and frameworks kids can use for the rest of their lives — including this one.