Homework scores went up 18%. Exam scores went down 20%. Same students, same six months, same AI tool. A working paper tracking 26,811 secondary students in China for 30 months just put a number on something teachers have suspected since ChatGPT showed up in every classroom: using AI to finish homework faster is not the same thing as learning the material — and the gap between the two shows up exactly where it matters most, on a test with the AI turned off.
This isn't the first study to find that pattern — a smaller 2024 RCT with GPT-4 tutors in Turkey found something structurally identical. But at 26,811 students and 30 months, this is by far the largest and longest-running evidence yet, and it's specific enough — down to which subjects, which students, and roughly what share of the effect is pure outsourcing — to actually change how you'd design an AI study habit, not just whether to worry.
Update — August 22, 2026: The Economist covered this same study (a 26,000+/27,000-student framing of the same DP21577 dataset, by the same Stockholm University / University of Hong Kong team), naming Doubao and DeepSeek as the two most-used tools among the roughly 80% of students who adopted AI. The coverage — and the Hacker News discussion around it — surfaced two sharper details we've folded in below: a 65-minute homework-time threshold that isolates exactly when the penalty kicks in, and a breakdown showing the effect is largest for previously high-achieving students, not smallest. See "The time-and-achievement breakdown" below.
TL;DR
| Question | Direct answer |
|---|---|
| What raised, what fell? | Homework scores +18%, homework completion time -30%, monthly exam scores -20% within 6 months, college entrance exam scores -18 to -24% |
| How many students, how long? | 26,811 students, grades 7-12, tracked for 30 months across a county in central China |
| Who ran the study? | David Strömberg, Victor Lei, and Yanhui Wu — Stockholm University and the University of Hong Kong — published via CEPR (DP21577) |
| What actually caused the exam drop? | ~80% of the loss is concentrated in students whose usage pattern matches "outsourcing" — very short homework time plus high homework scores |
| Does this mean all AI tutoring is bad? | No — guardrailed tools (hint-only prompts, answer grading instead of answer-giving) show the opposite effect in other studies |
| What should I actually do? | Use AI to check work and explain concepts after attempting the problem yourself, never to produce the answer you hand in |
What the study actually measured
The paper — "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education", CEPR Discussion Paper DP21577 — used 30 months of panel data on 26,811 students in grades 7 through 12 across a county in central China, tracking both homework performance and exam performance as generative AI tools got adopted at different rates across classrooms. That panel structure is what makes the finding stronger than a single before/after comparison: the researchers could watch the same students' homework and exam trajectories diverge over time as AI use increased, rather than just comparing two snapshots.
The headline numbers:
- Homework scores rose 18%, and completion time fell 30% — AI made homework faster and higher-scoring, exactly as you'd expect.
- Monthly exam scores fell 20% within six months of AI adoption — a test taken without AI access, on the same material.
- College entrance exam scores fell 18-24%, with the full size of the penalty only showing up after about two years — the highest-stakes exam in the dataset took the longest to reveal the damage.
- Losses were largest in social science subjects, then STEM, then languages, and were especially large for junior-grade students, high-achieving students, and boys.
The mechanism: outsourcing, not AI itself
The most important finding isn't the headline gap — it's what's driving it. The researchers found that roughly 80% of the exam-score decline was concentrated in students whose behavior pattern matched homework outsourcing: exceptionally short homework completion time combined with unusually high homework scores. That combination is the signature of a student pasting a question into a chatbot and copying the answer back, rather than attempting the problem and using AI to check or explain it.
That distinction matters because it means the "penalty" isn't a property of generative AI as a technology — it's a property of how it gets used. A student who spends the same amount of time on a problem, but uses AI to get unstuck or verify their reasoning, doesn't show the same signature in the data as one who outsources the whole task. The researchers' own framing, echoed across the coverage: for students, completing homework efficiently was never the goal — learning from it was, and generative AI made it trivially easy to optimize for the wrong variable.
This is the same mechanism explainx.ai has already covered on the other side of the age range: our AI-driven de-skilling piece covers an Anthropic RCT finding a 17% comprehension deficit in professional developers who leaned on AI coding assistants instead of writing code themselves. Same shape of result, same underlying cause, different population — practice that looks productive in the moment quietly erodes the skill it was supposed to build.
The time-and-achievement breakdown: what drives the 80%
Two details surfaced in The Economist's coverage and a widely-cited Hacker News breakdown of the underlying paper (credit to commenter kzz102, who quoted the paper's own excerpts) sharpen the "outsourcing, not AI" explanation into something closer to a mechanism:
- A 65-minute cutoff separates the two groups almost exactly. Students who still spent more than 65 minutes on homework — regardless of whether they used AI — scored about the same as non-AI students. But that group wasn't representative of long-term AI users: it consisted entirely of students who had adopted AI five months or less earlier. By six months post-adoption, no AI-using student in the dataset was spending more than 65 minutes on homework at all. The habit of finishing fast forms fast, and once it forms, the high-effort tier disappears.
- In the 50-65 minute overlap band, AI and non-AI students post statistically similar exam scores. This is the detail that rules out the simplest alternative explanation — that AI-assisted studying is just inherently worse, minute for minute. It isn't. Students who used AI but still put in comparable time did about as well as students who didn't use it at all. The penalty tracks time displaced from practice, not the quality of AI-assisted practice itself.
- The penalty is largest for previously high-achieving students, not smallest. Splitting students into achievement terciles based on prior performance, the exam-score effect ranges from -16% for the bottom tercile to -24% for the top tercile — a 50% larger hit for students who were doing best before AI showed up. One plausible read raised in the same discussion: high prior achievement is partly a proxy for willingness to grind through hard problems, and AI access is precisely what erodes the incentive to keep doing that.
Put together, these three findings turn the outsourcing story from a correlation into something closer to a dose-response curve: less time on the problem, less learning, and the students with the most room to lose that habit (previously high performers) lose the most.
Guardrails change the outcome — this isn't the first study to show it
This isn't a lone data point, and it isn't proof that AI-assisted learning is doomed. A 2024 randomized controlled trial with nearly 1,000 Turkish high school math students — run by Hamsa Bastani, Osbert Bastani, and colleagues, and later published in PNAS — split students into three groups: unrestricted GPT-4 access, a "GPT Tutor" version that gave hints instead of direct answers, and a no-AI control. The result was almost a preview of the Chinese study's finding, compressed into one semester: unrestricted GPT-4 access boosted in-session practice scores 48%, but once access was removed, those same students scored 17% worse than the control group on an unassisted test. The hint-only GPT Tutor group, given access to the identical underlying model, showed the improvement without the same collapse when AI was removed — the guardrail, not the model, was the deciding factor.
explainx.ai's own coverage of Dartmouth's Phosphor study points the same direction from a different angle: a platform that graded written answers against rubrics instead of supplying them was linked to a 0.71-1.30 standard deviation final exam gain, with 90.2% voluntary adoption — while the platform's actual open-ended chatbot barely got used at all. The lesson across all three studies is consistent: AI that grades, hints, or quizzes tends to help; AI that answers tends to hurt, and the difference is entirely in the guardrails, not the underlying model capability.
What this means if you're using AI to study
The practical takeaway isn't "stop using AI for homework" — it's stop using it as an answer machine and start using it as a tutor. explainx.ai's family framework for AI and homework puts this as a single line simple enough for a ten-year-old to apply: AI can teach you. AI cannot be you. A tutor explains a concept differently, checks your work, and quizzes you. A ghostwriter produces the assignment while you watch. The Chinese study's 80%-outsourcing finding is empirical backing for exactly that distinction — the students who used AI like a ghostwriter are the ones who lost ground; the ones who didn't largely didn't.
A concrete self-check that follows from the guardrail research above: the explain-it-back test. If you can't explain your own answer with the AI tool closed, you haven't learned the material — you've borrowed a correct-looking output that will not be there for you on the exam. That's precisely the mechanism the study's "outsourcing signature" (short time, high score) is detecting at scale.
This is also the design principle behind Melo, explainx.ai's own AI learning copilot: its Quiz, Practice, and Explain Back modes are built to make you produce the answer and defend it, the same shape as Phosphor's constructed-response questions and the Turkish study's hint-only GPT Tutor, rather than a chat window that hands you a finished answer to paste. If your current AI habit is "ask, copy, submit," swapping in a Quiz or Explain Back pass on the same material is a low-effort way to get the homework speed-up without the exam-score penalty this study documents.
The 65-minute finding above is exactly why that design choice matters more than it might sound. The penalty isn't about AI-assisted minutes teaching less per minute — it's about AI erasing the minutes altogether once outsourcing becomes habit. A tool that hands you a finished answer removes time from the ledger; a tool that makes you attempt the answer first and defend it after keeps the clock running on the part of homework that was actually doing the teaching. That's the whole design bet behind Melo's Quiz and Explain Back modes: preserve the time-on-task the CEPR data shows is the actual variable at stake, not just the appearance of homework getting "done." Our interactive learning pathways apply the same principle at the course level — structured practice with checks, not just a chat box.
What people are asking
"Is this specific to China's education system?" The mechanism isn't — homework-outsourcing behavior and the practice/test gap it creates aren't unique to any one country's curriculum. What's specific to this study is its scale and duration (26,811 students, 30 months), which is why it's useful as the largest data point we have, not the only one — it lines up with the smaller Turkey RCT and with Dartmouth's guardrailed-tool result.
"Should schools just ban AI for homework then?" The study doesn't test that policy directly, and the guardrail research above suggests banning AI outright forfeits the 18% homework-score and 30% time-savings benefit for no reason — a hint-only or answer-checking tool captured the upside in the Turkey RCT without the downside. The more defensible policy, backed by this data, is restricting what kind of AI access students get (hints and checking vs. direct answers), not whether they get any at all.
"How long before the damage shows up?" Fast on lower-stakes assessments — monthly exam scores dropped 20% within six months — but the full size of the effect on the highest-stakes exam (college entrance) only appeared after about two years. That lag is itself a warning: a student (or a school) could look at six months of rising homework scores and falling AI dependency-awareness and conclude things are fine, well before the real cost shows up on the exam that counts most.
Related on explainx.ai
- The OECD "28-Point AI Penalty" Claim, Fact-Checked — a 760,000-student OECD PISA report finds the same shortcut-vs-tutor split at global scale
- Dartmouth's Phosphor study: what a guardrailed AI tutor actually did
- The research behind AI-graded quizzes: 20 studies on interactive textbooks
- AI and homework: house rules that actually work
- AI-driven de-skilling: the developer version of the same problem
- Introducing Melo: explainx.ai's AI learning copilot
- Melo's generative UI for real-time learning
- Introducing interactive AI learning pathways
- ChatGPT for teens: safety features and Study Mode
Official sources: CEPR Discussion Paper DP21577 · SSRN preprint · Bastani et al., "Generative AI Without Guardrails Can Harm Learning," PNAS · Hacker News discussion of the Economist coverage
Figures in this post reflect the CEPR working paper and PNAS study as published; working papers can be revised before final peer-reviewed publication, and effect sizes from observational and quasi-experimental data should be read as strong evidence, not absolute certainty.
