A new study says something many daily users of chatbots half suspect: a short stretch of getting instant answers makes it harder to stick with a hard problem afterwards. In three randomized controlled trials with 1,222 participants, people who had an AI assistant for about 10 minutes did better while the tool was on, then gave up more and got fewer problems right once it was taken away. The paper, "AI Assistance Reduces Persistence and Hurts Independent Performance," is on arXiv, and UC Berkeley summarized it this week.
It sits next to a growing pile of explainx.ai coverage on the cost of leaning on models, including our look at cognitive debt when you stop retyping LLM code and the study on AI advice and cognitive surrender. This one is different because it measures behavior directly, with a control group, and it measures it fast.
TL;DR: the study at a glance
| Question | Answer |
|---|---|
| Who wrote it? | Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey (per the arXiv listing) |
| How many people? | 1,222 across three randomized trials (354, 667 and 201) |
| What tasks? | Basic fraction problems and SAT-style reading comprehension |
| What was the AI condition? | ChatGPT open alongside the task, answers allowed |
| When was it removed? | After 12 of 15 problems in the fraction experiments |
| Main result | AI group started more accurate, then solve rate fell and skip rate rose after removal; the no-AI group held steady |
| Time scale | About 10 minutes of interaction |
| Biggest caveat | Short tasks, online participants, no non-AI helper control as far as the summary describes |
What exactly did the researchers do?
According to Berkeley's write-up, the first two experiments gave participants 15 basic fraction problems. A control group worked unaided. The treatment group had ChatGPT available and could ask it for help, including for the answer outright. After the 12th problem, the researchers removed the AI and watched the last few problems.
Experiment 1 had 354 participants and Experiment 2 had 667, a larger replication. Experiment 3 had 201 participants and switched to SAT reading comprehension, to check that the effect was not specific to arithmetic.
Two measures matter here. The solve rate is plain accuracy. The skip rate is how often someone gave up on a question, which is the paper's proxy for persistence. The skip rate is the interesting one: it captures not whether you can do the task, but whether you are willing to keep trying.
Mapping the way through a maze, an image for how AI assistance reduces persistence on hard problems
What did they find?
The AI group began ahead. That is unsurprising, and it matters: the assistant genuinely helped. The change came when the help disappeared. Berkeley reports a sharp drop in the AI group's solve rate and a sharp rise in its skip rate after removal, while the no-AI group's solve rate held steady or rose and its skip rate stayed flat. In Experiment 3, both persistence and accuracy fell after the AI was withdrawn.
The Berkeley article does not give exact effect sizes or p-values, so we are not quoting any. The paper is the place to check those, and we have not independently reproduced the analysis.
The authors' account, in the arXiv abstract, is that current AI is optimized for instant, complete answers. That teaches people to expect them, and it denies people practice working through challenges. They write that persistence is foundational to skill acquisition and argue for AI that supports long-term competence alongside immediate task completion.
What did the authors say?
Brian Christian, a research fellow at Berkeley's Center for Human-Compatible AI and a co-author, told Berkeley that relying on an AI tool for just 10 minutes "significantly impairs our ability to focus and persist with difficult tasks." He added: "This is not a story about fractions and SAT problems."
He also argued the design of assistants is part of the problem. The tools, he said, "are often helping us in ways that are kind of unhelpful." His suggested direction is responses that behave more like tutors, guiding people instead of simply handing over the answer, and he warned that overreliance could eventually harm scientific expertise, such as staying close to the data.
Why the finding is plausible
The mechanism is familiar to anyone who has learned a hard skill. Struggle is where the learning happens, and effort tolerance is itself a trainable habit. An assistant that answers instantly resets your sense of how long a problem should take. When the answer no longer arrives in two seconds, a ten-minute grind feels unreasonable, so you skip.
We saw related effects in other studies we covered. The Dartmouth AI tutor study found large gains when the tutor was built to guide rather than tell, and the piece on homework scores versus exam scores described students whose graded work rose while their unaided exam results suffered. Read together, the pattern is that the same model can help or hurt depending on whether it substitutes for the thinking or scaffolds it.
What are people saying on Hacker News?
The Hacker News thread had 44 points and about 55 comments when we read it, and the reaction split. Treat these as the commenters' own reports and opinions, not findings.
- User bfung called it a weak study and wanted a control such as a calculator, including one that sometimes gives wrong answers, expecting similar persistence effects. imeanwhatdoikno made a related point about adding groups with books, calculators or human experts.
- dcanelhas described a pattern from coding: leaning on an assistant until it hits a wall, then paying the debt with interest when you must work it out yourself. They said it is not permanent and needs a recovery period.
- contubernio reported the opposite: using AI in mathematics research for about six weeks left them more productive and motivated.
- wccrawford compared it to a table saw making hand sawing feel unreasonable, arguing that tools change which effort is worth it.
The calculator and tools critique is the strongest one. A fair reading is that the study shows what happens when a very capable assistant is removed mid-task, which resembles real life when a model is down, rate-limited, or wrong, but does not by itself say whether the trade is net good.
What does the study not show?
Be careful with the headline. A few limits are worth stating plainly.
First, the tasks were short and low-stakes, and the participants were recruited online. A person doing paid, meaningful work may behave differently. Second, the measurement is immediate, so it says nothing about whether persistence recovers after a night's sleep or a week off. Third, the AI condition let people ask for answers directly, which is the worst case for learning. A hint-only assistant is a different treatment, and the authors themselves point toward that as the fix.
None of this makes the finding wrong. It makes it a well-powered signal about one specific usage pattern: answer-on-demand, then removal.
A stepwise route of stepping stones, an image for tutor-style AI that guides instead of answering
What should builders and heavy users do?
If you build products with models, the paper is an argument for a learning mode. Our write-ups on AI tutors that teach by making and on ChatGPT study mode for teens show vendors already moving toward guided responses. Practical defaults that fit the evidence:
- Hint ladders. Offer a nudge first, a worked step second, the full answer last.
- Ask the user to attempt first. Even a one-line guess changes the interaction from consuming to checking.
- Fade support. Reduce assistance as a user improves instead of keeping it constant.
- Plan for outages. If your product is a skill-building tool, test what users can do when the model is unavailable.
If you are a heavy user, the cheap habits are to draft your own first attempt on anything you want to keep being good at, to ask for questions or critique rather than finished output, and to keep a small amount of AI-free practice. For code, retyping or explaining generated code to yourself is a low-cost version of the same idea.
A nearly finished puzzle with one piece left, an image for practising hard problems without AI help
Bottom line
The study does not show that AI makes you worse at thinking. It shows that ten minutes of answer-on-demand help can make the next hard problem feel less worth sticking with, and that this showed up across 1,222 people, two task types and three trials. The more useful reading is about design: the same model can erode or build persistence depending on whether it hands over answers or guides you toward them. Expect the next round of research, and product features, to test exactly that.
Related reading
- Cognitive debt: retyping LLM code
- AI advice and cognitive surrender study
- Dartmouth AI tutor study and effect size
- AI homework scores and the exam learning penalty
- AI tutors and learning by making
- ChatGPT study mode and teen safety
- Primary sources: arXiv paper and Berkeley research news
Details reflect the arXiv listing (version 5, October 3, 2026) and Berkeley's write-up as of October 11, 2026.
