Congress Couldn’t Write AI Rules for Schools. 98 Teenagers Did.
There's no national policy for how AI should be used in schools. Not one. States are passing their own patchwork of rules, districts are writing their own, and most families…
There’s no national policy for how AI should be used in schools. Not one. States are passing their own patchwork of rules, districts are writing their own, and most families are left guessing.
So a group of teenagers did what Congress hasn’t: they wrote the bill themselves.
On a late-July weekend, 98 teens from all 50 states gathered at the Edward M. Kennedy Institute for the United States Senate — inside a full-scale replica of the actual Senate chamber — and spent the weekend debating, amending, and voting on legislation to govern AI in K-12 schools. They called it S 2026, the Students First Act. It passed 83 to 15.
Here’s the part worth sitting with: the student-focused section of that bill has 15 provisions, and several of them circle back to the same worry, again and again, that AI is doing their thinking for them.
That’s not a rule an ed-tech company would write. It’s not a rule most school boards have written either. It’s a rule from the people actually sitting in the classroom, who can apparently tell the difference between AI that helps them think and AI that thinks for them, even when the adults writing policy around them mostly can’t.
Why This Matters More Than Another AI-in-Schools Headline
At Auxesis, we’ve spent a long time naming this exact problem: the illusion of mastery. It’s what happens when a student looks like they’ve learned something, the essay reads fine, the answer is correct, but the actual thinking never happened. Good output, no understanding underneath it.
Usually we’re the ones making that case, citing a study or a researcher. This time, 98 teenagers made it themselves, from the inside, without anyone telling them to. They didn’t call it the illusion of mastery. They called it AI doing their thinking for them. Same problem, plainer words.
That cuts against the two lazy stories adults tend to tell about kids and AI: either kids will use it responsibly with no guardrails needed, or kids can’t be trusted with it at all. The teens at the Kennedy Institute didn’t take either position. They drew a real line, AI for editing and brainstorming, yes, AI for the writing itself, no, the same kind of line Auxesis draws with the CPA framework (Concrete, Pictorial, Abstract) in math: AI can support you at any stage, but it can’t do the stage for you. Skip the struggle, and the mastery you end up with is fake, no matter how correct the final answer looks.
One of the provisions the teens landed on: no AI for writing. Editing, brainstorming, and studying with AI would be allowed starting in eighth grade — but the writing itself has to stay theirs.
What This Means for You
If you’re a homeschool parent, this is permission to trust your instinct that letting AI just explain it and letting AI help you think it through are not the same thing, a group of teenagers just spent a weekend proving they know that too.
If you’re an educator, this is a template. The Students First Act draws a line by task, not by blanket ban: editing and brainstorming are fine, the core cognitive work isn’t. That’s a more useful starting point for your own classroom AI policy than most of what’s floated around this year, and it’s the same logic behind Auxesis’s Facilitator Certification, which is built around exactly this judgment call: knowing when AI belongs in the room and when it doesn’t.
There’s still no national policy. But for one weekend, in a replica Senate chamber in Boston, 98 teenagers showed they’d already figured out what the adults are still arguing about.
Source: NPR, July 30, 2026 (npr.org/2026/07/30/nx-s1-5853571/students-set-ai-policy)
Team the 98 teens’ rule, AI for editing and brainstorming, never for the writing itself, or team stricter: no AI at all until high school? Curious which side your house or classroom lands on.
Want a framework for making this call yourself? The Parent Track and Educator Track both walk through exactly where AI belongs, and where it doesn’t.
Two weeks ago, a new AI tutoring company called Bloomy launched out of Y Combinator with a number attached: students in an early pilot grew 1.8 times faster than expected on the NWEA MAP assessment, a widely used test of academic growth. The founder has been careful — he’s called it an observational pilot, not a randomized study, and the sample was about 150 middle schoolers at one Massachusetts charter school.
That caveat matters, and other outlets have already covered it well: no control group, no randomization, one school. If you want the stats critique, it’s out there.
We want to talk about something the coverage has missed — because it’s sitting in Bloomy’s own product design, not in a footnote.
The part of Bloomy that looks a lot like us
Bloomy’s platform moves students through three stages for every skill: Base Camp (worked examples), Climb (guided practice with an AI tutor), and Summit — an unaided, ten-question test with the tutor deliberately absent. Students don’t advance until they clear roughly 90% on Summit.
If that structure sounds familiar, it should. It’s the same logic behind our Mastery Gates: Explain, Apply, Sustain. A student doesn’t move forward because they got through the material. They move forward because they can perform without support, cold, when nobody’s coaching them through it.
Here’s the question that matters: why build that gate at all?
Because a growth number doesn’t tell you what you think it tells you
NWEA MAP growth is a well-respected, widely used measure — but it measures performance on a norm-referenced test. It tells you a student answered more items correctly, or harder items correctly, than before. It does not, by itself, tell you why — whether a student built durable understanding, or got faster at pattern-matching the kind of problems that show up on that kind of test
That gap between “the score went up” and “the understanding is real” is what we call the illusion of mastery: the condition where a student looks like they’ve learned something, but the underlying thinking never actually happened. Good grades without reasoning. A rising RIT score without a student who can explain what they did.
The interesting thing is that Bloomy’s own architecture seems to already know this. If a rising test score were sufficient proof of learning, you wouldn’t need a Summit gate. You’d just assign more problems and let the score climb. The fact that Bloomy built an unaided, high-bar checkpoint before letting a student move on suggests someone on that team understands the difference between “the number went up” and “the student can actually do it alone” — even while the marketing leans on the number.
What this means for you
If you’re evaluating any AI tutoring tool — Bloomy or otherwise — for your family or your classroom, the growth-percentage headline is the least useful thing to ask about. The better questions:
Does the tool require a student to perform without help before it lets them move forward?
Can you see why a student got something wrong — a specific misconception — not just that they got it wrong?
Does “mastery” mean “answered correctly,” or does it mean “can explain the reasoning, apply it somewhere new, and still have it three weeks later”?
That’s the standard we hold our own Mastery Gates to, and it’s the standard worth holding any tool to — including ours.
If you’ve spent any time in parenting or education circles this year, you’ve seen the headlines: AI is making kids worse at thinking. Teachers can’t get students to reason anymore. There’s a new term for it – “the great unwiring” – and the evidence being passed around usually points to one study: nearly 1,000 high school math students in Turkey, split into groups, given access to ChatGPT-4 while they studied. The unrestricted group did 48% better on practice problems. Then, when the AI was taken away for the final exam, that same group scored 17% worse than students who’d never used AI at all.
That statistic is real. It comes from a peer-reviewed study published in PNAS – “Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics” (Bastani, Bastani, Sungu, Ge, Kabakci, and Mariman) – and it’s the study most people are pointing to when they cite this fear. It’s part of a broader wave of alarm: Fortune has covered a separate Brookings report warning of a “great unwiring” of students’ brains, and Axios has covered fresh polling showing most teachers think AI is hurting kids’ critical thinking. Every version of the story – whichever one you read – stops at the scary number.
Almost none of them mention the group that changes the whole story.
The Detail Getting Cut From Every Retelling
The researchers didn’t test one AI condition. They tested three: a control group with only textbooks and notes, a group with unrestricted ChatGPT-4 access (“GPT Base”), and a third group using a version of GPT-4 built with teacher-designed guardrails that gave hints and guiding questions instead of direct answers (“GPT Tutor”).
The GPT Base group is the one making headlines – 48% better on practice, 17% worse on the final exam. Used as a crutch, then gone, and the students had nothing underneath them.
The GPT Tutor group told a different story entirely. On practice problems, they didn’t just match the Base group – they outperformed it by 127%. And when the AI was removed for the final exam? They scored about the same as the control group. No penalty. No collapse. The guardrails held.
Synthesis Tutor calls itself, right there in the page title, “the world’s first superhuman math tutor.” That’s not marketing shorthand somebody exaggerated in an ad — it’s the actual tagline on their website today.
It’s a big claim. So when a claim is that big, the useful question isn’t “does that sound impressive?” It’s “what’s the evidence, exactly, and does it hold up?”
Synthesis actually makes this easy. They wrote a blog post called “Does the Synthesis Tutor get results?” and it names their source directly: a DARPA-funded program called the Digital Tutor. Follow that thread, and here’s what you find.
What the DARPA study actually was
The Digital Tutor was real, and by the numbers reported, it was genuinely impressive — for what it was designed to do. DARPA wanted to see how close a piece of software could get new recruits to the expertise of a seasoned professional, fast. So they built a tutoring system and tested it on U.S. Navy sailors training to become Information Systems Technicians — the people who keep a ship’s IT systems running.
Over 16 weeks, those sailors went through the Digital Tutor program. At the end, they were tested against two other groups: Fleet technicians with an average of ten years on the job, and sailors trained the traditional classroom way. The Digital Tutor group outperformed both, by a wide margin, on troubleshooting real IT systems.
That’s a strong result. It’s also, if you look at the actual DTIC records, from assessments run in 2010 — sixteen years ago now — on adult Navy technicians learning enterprise IT systems. Not elementary students. Not math. Not children at all.
Where the gap is
None of this means Synthesis Tutor is bad, or that the company is lying. They’re transparent about their source — they link straight to the DARPA documentation and a public summary of it. That’s more citation than most ed-tech marketing bothers with.
But there’s a real distance between “a 2010 program that helped Navy sailors master IT troubleshooting faster than classroom training” and “the world’s first superhuman math tutor” for a 6-year-old learning to subtract. The first is a specific, well-documented result in a narrow, adult, technical-skills domain. The second is a sweeping claim about a completely different subject, age group, and learning context — resting on that same study as its evidence.
That gap is the actual lesson here, and it’s a useful one to practice noticing — for you and for your kids.
A framework for checking claims like this
This is exactly the kind of moment where the third piece of metacognition — Evaluate — earns its keep. Monitor asks “do I understand this?” Regulate asks “what should I do differently?” Evaluate asks the question most of us skip: “was that reasoning actually sound?”
Applied to a marketing claim, Evaluate looks like three quick questions:
What’s the actual claim? (“Superhuman” — a specific comparative claim, not just enthusiasm.)
What’s the cited evidence? (A study — good, that’s better than nothing.)
Does the evidence match the population and the product? (Adult Navy IT trainees, 2010, technical troubleshooting — vs. a math app for a 7-year-old, in 2026.)
That third question is where most impressive-sounding claims quietly fall apart. It’s not about catching companies in lies — it’s about training the habit of checking whether the receipt matches the bill. That’s a skill worth modeling out loud with your kids the next time an app, a toy, or a tutor promises something “revolutionary.”
The takeaway
Citing a source is good. Citing the right source is what actually matters. Before any tool — ours included — earns a claim like “this works,” it should be able to show you evidence that was actually about the kids, the subject, and the context you’re deciding for. If it can’t, that’s worth knowing before you buy in.
Auxesis doesn’t sell superhuman anything. We teach the slower thing that actually works: struggle, checked understanding, and frameworks kids can use for the rest of their lives — including this one.
If you run a homeschool co-op, a microschool, or lead professional development for other educators, here’s a piece of news worth five minutes of your attention.
The National Science Foundation just awarded $11 million to the Computer Science Teachers Association to run something called AI Professional Development Weeks — a multistate effort to train K-12 teachers in AI and computer science fundamentals. Over the next two years, it will run in Indiana, South Carolina, Minnesota, New Jersey, Iowa, and Illinois, plus at least three more states still being added. The program will directly train roughly 2,500 to 3,000 teachers, with a goal of reaching 500,000 to 600,000 students through them.
That’s a real, federally-funded bet that teachers — not just students — need structured AI training. It’s also, if you lead a co-op or a microschool, a signal worth reading closely, even though almost none of the coverage so far has said a word about you.
What the coverage is missing
Every article on this program so far has been written for a district administrator or a policy reporter. It covers funding numbers, participating states, and rollout timelines. What it doesn’t cover – because it isn’t built to – is what any of this actually looks like once a teacher walks back into a classroom with a week of AI training under their belt.
A week of AI PD can teach a teacher what tools exist and how to use them. It can’t teach them how to use those tools without quietly letting AI do the students’ thinking for them – the difference between augmenting a lesson and replacing the learning inside it.
Why this matters for co-ops and microschools specifically
You likely can’t register for this one. The one piece of eligibility detail available – Illinois’s stipend criteria – limits funded participation to “public and public charter school teachers,” and nothing in CSTA’s public materials suggests a path in for homeschool co-ops or independent microschools. But the demand it’s responding to is one independent educators feel just as much, with no federal grant answering it.
Here’s the concrete version: imagine a co-op teacher who hears about AI PD Weeks, wants the same caliber of training, but isn’t in one of the funded states or isn’t part of a public district at all. She doesn’t need a district-scale program. She needs the same underlying question answered: how do I use AI in my teaching in a way that builds my students’ thinking instead of shortcutting it? That’s a pedagogy question, not a tools question – and it’s the one federal AI-in-education funding, so far, hasn’t been built to answer.
Where Auxesis fits
This is the exact gap the Educator Track exists to close – not tool training, but the facilitation layer underneath it. Moving from “explainer” to “facilitator,” learning to ask instead of tell, and using COMPASS as the pedagogy layer that sits on top of whatever AI tools a teacher already has access to. A teacher doesn’t need to wait for a state grant to start building that skill.
The takeaway
$11 million in federal funding for teacher AI training is a genuine signal: this is now considered infrastructure, not an extra. If you lead a co-op, a microschool, or PD for other educators, the honest takeaway isn’t “we should get a grant like that.” It’s that the underlying need – teachers who can use AI well, not just teachers who’ve seen a demo – is universal, funded or not. The Educator Track was built to meet that need directly, no state eligibility required.
If your child’s homework scores have quietly climbed this year, that’s good news — unless AI is doing more of the thinking than your child is. A new study of 26,811 students in China just put hard numbers on a pattern a lot of parents have felt but couldn’t prove: AI can make homework look better while making the learning underneath it worse.
Researchers from Stockholm University and the University of Hong Kong tracked these students for 30 months — homework scores, completion time, monthly exams, and entrance-exam results, across nine subjects. Here’s what they found. When students started using generative AI on homework, their homework scores went up 18%, and they finished 30% faster. Sounds great. But within six months, their monthly exam scores — taken closed-book, no AI allowed — dropped 20%. And on the exams that matter most, the high-stakes entrance exams, scores fell 18% to 24%, with the full damage only showing up after about two years.
The researchers found the losses weren’t spread evenly. About 80% of the AI-using students showed a specific pattern: exceptionally short homework time paired with unusually high homework scores. In plain terms — they weren’t using AI to check their work or get unstuck. They were handing the thinking over entirely and turning in the result. The study calls this “homework outsourcing,” and it’s the group where almost all of the exam damage showed up.
Worth saying plainly: this is one study of Chinese secondary students, grades 7 through 12, not U.S. homeschoolers. The exact numbers won’t map onto your household. But the mechanism the study describes — a tool that produces a correct-looking answer without requiring the struggle that builds understanding — isn’t specific to China, or to any curriculum. It’s specific to what AI homework tools do.
Why this happens: the gap between Explain and Sustain
At Auxesis, we talk about three conditions a student has to clear before we call something “mastered”: they can Explain it in their own words, Apply it to a new problem, and Sustain it — meaning it still holds up weeks later, without a crutch in the room.
A homework score only tests the first condition, and sometimes not even that — it tests whether a correct answer got produced. AI is extremely good at producing correct answers. It is not able to build the second or third condition for your child; that only happens through the struggle of getting an answer wrong, figuring out why, and trying again. A rising homework grade with a sinking exam grade is what it looks like when a student is clearing gate one on borrowed thinking and never reaching gates two and three at all.
What this looks like at your kitchen table
Say your child brings home a worksheet on solving for x, and it’s done in ten minutes with every answer correct. That used to be a great sign. Now it’s worth one follow-up question before you sign off on it: “Can you solve one more like this — right now, out loud, no notes?” If they can walk you through it cleanly, the homework score was earned. If they stall or reach for a device, the homework score was borrowed, and today’s a good day to slow down and actually work one problem together.
Three guardrails, not a ban
You don’t need to take AI away to fix this — the study’s own data suggests the damage comes from a specific misuse pattern, not from AI existing in the house.
Closed-book check-ins. Once a week, pick one homework problem your child already turned in and have them redo it with no tool in hand. This is the fastest way to see whether a grade reflects understanding or output.
Subject-specific limits, not blanket ones. The study found losses were worst in social science subjects, then STEM, then languages. If you only have bandwidth to watch one subject closely, watch writing and reasoning-heavy work before you watch math drills.
Ask what mode it’s in. Some AI tools default to giving answers; some can be pushed into asking questions instead (the difference between an answer machine and a Socratic tutor). Which mode your child’s tool is in matters more than whether they’re allowed to open it.
None of this requires panic. It requires knowing the difference between a grade and an understanding — and checking for the second one every so often, out loud, without the tool in the room.
If you want a simple way to build that check-in into your week without becoming the homework police, the Parent Track walks through exactly this kind of mastery tracking — a five-minute weekly habit, not a new curriculum.
Half of U.S. school districts now say they’ve trained their teachers on AI. A year earlier, it was about a quarter. That’s a real, fast shift — backed by a nationally representative RAND survey of the American School District Panel, not a press release.
Here’s the part worth sitting with: when RAND’s researchers asked district leaders what “training” actually meant in practice, the honest answer, over and over, was we made it up as we went. Eleven of the fourteen district leaders RAND interviewed built their AI training programs themselves, from scratch, because the outside options weren’t good enough. One leader put it plainly: “There are people that are claiming to have the best practices and are making money hand over fist… if they claim to be telling you best practices, they don’t have them yet. They don’t exist yet.”
That’s not a knock on those district leaders — they’re doing real work under real time pressure. It’s a diagnosis of the moment. AI arrived in classrooms faster than anyone built the professional development to match it, so schools are running the experiment live, with actual students, mostly measuring success by whether teachers stopped being afraid of the tool.
That’s the real gap this piece is about. Most of the training districts describe is tool training: how ChatGPT works, how to write a prompt, how to use an AI lesson-planning assistant. That’s a reasonable first step — teachers who are anxious about a tool can’t use it well, and RAND found addressing that fear was nearly every district’s starting point. But tool literacy and facilitation are two different skills, and only one of them protects the thinking in the room.
A teacher who knows how to prompt ChatGPT can still end up, without meaning to, running a classroom where the AI does the reasoning and the student does the copying. Nothing about “how to use the tool” teaches a teacher to notice that moment, or to redirect it. That’s a facilitation skill — the ability to read whether a student is genuinely stuck or just handing off the thinking, and to ask the next question instead of supplying the next answer. It’s the same shift Auxesis’s Educator Track builds toward: moving from Explainer, who fills every silence with an answer, to Facilitator — from Practitioner to Coach.
Picture two versions of the same fifth-grade classroom, six months into AI rollout (illustrative, not a real case). In the first, teachers got a solid half-day on an AI planning tool and were told to “play around with it.” Word-problem homework improves fast — suspiciously fast. Kids paste the problem into a chatbot, copy the steps, turn it in. The teacher, trained on the tool but not on what to watch for, sees clean homework and reasonably assumes the class is ahead of schedule. In the second, teachers got the same tool training, plus one more habit: asking “walk me back through how you got there” before accepting an answer as done. Same tool. Very different classroom, six months in.
That second habit is what Auxesis calls a Facilitation Brief — a short, structured read before a session on where a student actually stands, so class time goes to facilitating instead of re-diagnosing from scratch. It’s one piece of COMPASS, the operating layer Auxesis builds around AI in the classroom: AI helps before the session and after it, in the briefing and the five-minute note; the actual teaching stays human, on purpose.
There’s also an equity story here, and it’s worth naming plainly. RAND found low-poverty districts have consistently trained teachers on AI faster than higher-poverty ones — 43% versus 6% in fall 2023, 67% versus 39% by fall 2024 — and district leaders’ own projections show that gap holding into the 2025–2026 school year, with almost all low-poverty districts trained and only around six in ten high-poverty districts there yet. Whatever training model turns out to work will reach wealthier schools first. That’s one more reason the training that does exist should be built around facilitation, not just tool onboarding — a district that gets one real shot at AI professional development shouldn’t spend it on prompt-writing alone.
None of this argues against training teachers on AI faster. It argues for training them on the right thing. The tool is the easy part to teach. Reading a classroom, protecting productive struggle, and knowing when to step back instead of stepping in — that’s the harder, more durable skill, and it’s the one most current AI-in-schools training is skipping past.
If you’re building or choosing AI professional development for your school or district, the question worth asking isn’t “does this cover the tool.” It’s “does this teach my teachers to notice when a student stopped thinking.” The Educator Track’s facilitation modules exist to answer exactly that — not as a replacement for tool training, but as the layer most programs are currently missing.
A follow-up to The 17% Tax — same research family, different question.
We wrote recently about the OECD’s finding that AI-assisted math practice can look great and still leave nothing behind — up to 17% worse performance once the AI is taken away. That post ended with one diagnostic question: “walk me through how you got this.”
There’s a second piece of evidence worth its own post, because it answers a question the math study doesn’t: how fast does the gap open, and does it show up outside of math?
The one-hour test
A study cited in the same OECD Digital Education Outlook 2026 had students across several US universities write a short essay — one group alone, one with a search engine, one with a general-purpose chatbot doing much of the drafting. One hour later, researchers asked each student to quote a sentence from what they’d just “written.” Among the unaided and search-engine students, 89% could. Among the chatbot group, only 12% could.
Worth being precise here, in the same spirit as the caveat we ran last time: the specific study behind this appears to be MIT Media Lab’s “Your Brain on ChatGPT” research (Kosmyna et al.) — a small trial (54 participants), still a preprint, not yet peer-reviewed. Different write-ups of it report slightly different numbers (some cite 90%/17% instead of 89%/12%), which is normal for early-stage research moving through secondary coverage, but it means this shouldn’t be treated as a settled, precise figure — just a strong, repeatable signal in the same direction as the math result: fast AI produces work that looks finished but was never really held by the student who “wrote” it.
One hour. Not a semester, not a unit test — sixty minutes was enough for four out of five students to lose their grip on their own sentences.
Why one question isn’t enough for this one
The “walk me through how you got this” check works well for a worked math problem because there’s a step-by-step path to retrace. Writing doesn’t hand you that same rope. A finished essay doesn’t show its work the way a solved equation does — which means the single-question check from the math post genuinely won’t catch this failure mode. You need something with more structure.
That’s what Auxesis’s Metacognition framework is for — three questions, asked in sequence, that work regardless of subject:
Monitor — “Before you turn this in: what’s the one sentence in here that’s most you? Point to it.” If your child can’t find one, that’s the tell — not a bad grade, just a flag that something outside their own head produced the words.
Regulate — “What would you do differently if you had to write this again without any help?” This isn’t a punishment question. It’s the moment that turns a flagged gap into an actual second pass — the regulation step is what separates “I noticed this wasn’t mine” from “I fixed it.”
Evaluate — Circle back a day or two later, unannounced, with a version of the one-hour test: “Explain the argument you made in that essay.” If they can’t, you’ve learned something real about the assignment — and caught it while it’s still cheap to address, not at the next test.
Why this can’t be a one-time fix
The math study and the essay study point at the same underlying mechanism from two directions: performance during the task tells you almost nothing about what’s going to stick. The only way to know is to check after — which is exactly what Monitor/Regulate/Evaluate is built to do, and exactly what a finished worksheet or a polished essay can’t tell you on its own.
This is also why it’s a framework and not a one-off question: the math post’s single check catches one failure mode; a repeatable three-step loop catches it across subjects, because “did my child actually think this through” isn’t a math-specific problem — it’s the same one whether the tool wrote an equation or a topic sentence.
If you want the full walkthrough of how we teach Socratic questioning at home — the same instinct that powers Monitor/Regulate/Evaluate — that’s covered in the Parent Track.
If you’ve used AI to help your child with homework, you’ve probably felt a small flicker of guilt about it. New research says that flicker is worth paying attention to — but not for the reason the headlines suggest.
Over the past month, a wave of studies has landed with real numbers attached to something a lot of parents have sensed intuitively: leaning on AI to get answers is different from using it to get smarter.
What the research found
A study of 1,222 people — conducted across teams at Oxford, MIT, UCLA, and Carnegie Mellon and reported by Forbes on June 30, 2026 — found that AI assistance measurably reduces independent performance and persistence, and that this effect shows up fast: within 10 to 15 minutes of AI-assisted work.
Separately, an analysis covering 26,000 students (reported by Psychology Today) found learning losses of 25 to 30 percent on math work once students started leaning on AI to get through it — the study tracked how much time students spent on problems that were easy to hand off to a chatbot, and found that less time on the problem meant less learning from it.
And in higher education, 90 percent of faculty surveyed now say AI is weakening students’ critical thinking.
Three different research teams, three different populations, one consistent finding: when AI does the thinking, the student doesn’t.
Why this isn’t actually surprising
None of this means AI is bad for learning. It means answer-first AI is bad for learning — which is a much more useful thing to know, because it tells you exactly what to change.
Struggle is not a bug in the learning process. It’s the mechanism. When your child sits with a hard problem, tries something, gets it wrong, and tries again, that friction is where the actual wiring happens. An AI tool that skips straight to the answer isn’t saving your child time — it’s skipping the part of the assignment that was actually the assignment.
This is the same idea behind metacognition — one of the core frameworks we teach in the Singapore Stack: the habit of monitoring your own understanding ("do I actually get this, or does it just look familiar?"), regulating your approach when you’re stuck ("what should I try differently?"), and evaluating afterward ("did that actually work, and why?"). A student who hands a problem to AI the moment it gets hard never runs this loop. A student who’s taught to run it — with or without AI in the room — builds the skill that the loop itself was practicing.
A concrete example
Say your child is stuck on a word problem. The answer-first move is: type it into a chatbot, get the answer, copy it down, move on. Ten seconds, zero learning.
The metacognitive move looks different: ask your child to explain what the problem is actually asking before touching any tool. Have them guess roughly what a reasonable answer would look like. Let them try it — get it wrong, even. Then bring in AI, not to hand over the answer, but to ask a question back: "What’s one thing you could check about your work?" or "What would happen if you tried it a different way?"
Same tool. Completely different result. One erodes the skill the homework was supposed to build. The other builds it.
What to do this week
You don’t need to become an AI expert to make this shift — you need about five minutes and a habit. Before your child uses any AI tool for schoolwork, run this quick check:
Has my child tried this on their own first, even briefly?
Is the AI being asked for an answer, or asked a question back?
Could my child explain how they got the answer, out loud, right now?
If the answer to #3 is no, the tool did the thinking instead of your child.
We’ve built this out into a fuller "5-Minute AI Use Audit" — a short, printable checklist you can keep next to the homework table. It’s free, and it’s below.
Why we’re writing this now
This isn’t abstract for us. Inside our own community, the single most-discussed question right now is simply: "How are you using AI with your kids?" Parents are asking this in real time, without a clear answer in front of them. This research — and the framework above — is our answer.
AI isn’t the enemy here. Skipping the struggle is. Used the right way, AI can be a genuine thinking partner for your child. Used the wrong way, it’s just a very fast way to look like you learned something you didn’t.
For two years, Auxesis has made an argument that felt, at times, like we were making it alone: that a student can look like they’ve mastered something — fast, fluent, high scores — while the actual thinking never happened. We call this the illusion of mastery. It’s the thing we exist to fight.
This month, two independent education writers walked into a well-funded, fast-growing math platform called Math Academy and came out describing the exact same problem, in almost the exact same words — without ever having heard of us.
What happened
Math Academy markets itself on speed. Math educator Michael Pershan spent a month inside the platform, working through a full course, to see what that speed actually produced. His account, Math Academy: A Mixed Review, is careful and fair, but one line does the most damage: “Math Academy offers direct instruction for procedures, discovery learning for concepts.” In practice, he found, the platform hands you the steps for solving a problem clearly and quickly — and leaves you to construct the why almost entirely on your own, with no one checking whether you actually did.
Pershan’s sharper point is about incentives, not just design. The platform awards experience points for completing exercises, he writes, not for reading the conceptual explanations sitting alongside them: “If the exercises don’t require the concepts, then the concepts only inhibit your progress and kids will drive past them at 75 mph.” A student can hit every target the app measures without ever slowing down for the part where understanding actually forms.
Dan Meyer picked up Pershan’s review and went after the marketing claim itself in a piece titled “It Is Fun to Pretend That Hard Things Are Easy!” He points out a pattern common to platforms that promise to have cracked the code on faster learning: they quietly redefine what “learning” means — not to educators, not to universities, not to the people who eventually have to use the math for something real, just to their own dashboard. Fast completion of exercises gets rebranded as mastery. The gap between the two is where the illusion lives.
Why this matters beyond one platform
Neither Pershan nor Meyer set out to make our argument. They set out to review a product. That’s what makes it useful. When a business built around a critique makes that critique, it’s easy to dismiss as self-interested. When two working educators, with no connection to each other or to us, independently spend real time inside a platform and land on the same conclusion — using almost the same language — that’s a signal the pattern is real, not a marketing angle.
The pattern, stated plainly, and stated by them, not us: speed and fluency are being measured and rewarded. Understanding is not. A system that only measures what’s easy to measure will quietly train students, and the adults watching them, to mistake the measurement for the thing itself.
What we’d ask instead
Auxesis builds around a different question: not “did the student finish,” but “can the student explain it, apply it somewhere new, and still have it a month from now?” We call that the difference between Explain, Apply, and Sustain — the three conditions that have to hold before we call something mastered. A student can pass a platform’s exercise set, by Pershan’s own account, and still fail all three. That gap is not a rounding error. It’s the whole problem.
None of this means adaptive practice tools are worthless — deliberate, well-designed practice has a real place, and Pershan says as much in his review. The problem isn’t that Math Academy exists. It’s that “fast” and “learned” are being treated as the same word, at a moment when a lot of anxious parents are looking for a shortcut and being told, credibly, that one exists.
The reviewers didn’t need us to tell them that. They found it themselves, inside the product, and wrote it down. Our job now is simple: point at what they found, and say it plainly, before “cognitive offloading” gets flattened into a buzzword nobody bothers to define.