If you’ve spent any time in parenting or education circles this year, you’ve seen the headlines: AI is making kids worse at thinking. Teachers can’t get students to reason anymore. There’s a new term for it – “the great unwiring” – and the evidence being passed around usually points to one study: nearly 1,000 high school math students in Turkey, split into groups, given access to ChatGPT-4 while they studied. The unrestricted group did 48% better on practice problems. Then, when the AI was taken away for the final exam, that same group scored 17% worse than students who’d never used AI at all.

That statistic is real. It comes from a peer-reviewed study published in PNAS – “Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics” (Bastani, Bastani, Sungu, Ge, Kabakci, and Mariman) – and it’s the study most people are pointing to when they cite this fear. It’s part of a broader wave of alarm: Fortune has covered a separate Brookings report warning of a “great unwiring” of students’ brains, and Axios has covered fresh polling showing most teachers think AI is hurting kids’ critical thinking. Every version of the story – whichever one you read – stops at the scary number.

Almost none of them mention the group that changes the whole story.

The Detail Getting Cut From Every Retelling

The researchers didn’t test one AI condition. They tested three: a control group with only textbooks and notes, a group with unrestricted ChatGPT-4 access (“GPT Base”), and a third group using a version of GPT-4 built with teacher-designed guardrails that gave hints and guiding questions instead of direct answers (“GPT Tutor”).

The GPT Base group is the one making headlines – 48% better on practice, 17% worse on the final exam. Used as a crutch, then gone, and the students had nothing underneath them.

The GPT Tutor group told a different story entirely. On practice problems, they didn’t just match the Base group – they outperformed it by 127%. And when the AI was removed for the final exam? They scored about the same as the control group. No penalty. No collapse. The guardrails held.

Same AI, opposite outcomes: GPT Base scored 17% below control on the final exam; GPT Tutor matched control

Leave a Reply

Your email address will not be published. Required fields are marked *