01
Why this, why now
Through 2025 and 2026 the evidence on AI and learning stopped being anecdotal and started being large, controlled and consistent. A run of careful studies now says the same thing from different directions: AI is neither good nor bad for learning. It amplifies whatever the human is asked to do with it.
We find that more useful than either the panic or the hype, because it makes the outcome a design choice rather than a property of the technology.
This edition changes two things. The newer literature refines the variable: it is not effort in general, it is a specific kind of cognitive engagement, and the distinction matters because it tells you which friction to keep. And it adds a section the July edition should have had. Every other dossier in this series carries a counter-argument that undercuts its own case. This one did not, and that was our oversight rather than our judgement, and the strongest version of that counter-argument is section 10.
Timeline
- 2023 · Ethan Mollick warns of a "Homework Apocalypse" as AI makes traditional homework unassessable.
- Jun 2025 · A Harvard randomised trial finds a well-built AI tutor beats active-learning classes.
- Jun 2025 · MIT's "Your Brain on ChatGPT" shows weaker neural engagement and poor recall in AI-assisted writers.
- Jul 2025 · Mollick publishes "Against Brain Damage", pushing back on how that study is being read.
- Mar 2026 · An NBER working paper takes up AI, human cognition and knowledge collapse.
- May 2026 · Anthropic publishes its own research on how AI assistance affects the formation of coding skills.
- 2026 · A 26,811-student study in China measures a learning penalty hiding behind higher homework grades.
- Jul 2026 · Nature Medicine names "never-skilling"; UChicago Law publishes an AI strategy built around AI-resilient assessment.
02
Contents
1. The paradox in two studies
2. The good case: a tutor that beats the classroom
3. The crutch: the grades went up, the learning went down
4. Cognitive debt: the brain on ChatGPT
5. What the newer work refines
6. Never-skilling: the skill that never forms
7. Engagement is the variable
8. It does not stop at school
9. The assessment problem
10. Where we could be wrong
11. Design for capability, not convenience
12. Your own people: reskill or de-skill
13. What we are watching
14. Verification and sources
03
1. The paradox in two studies
Two of the most careful studies of AI and learning published in the last two years reached opposite conclusions. Both are right, and we think reading them together is the only way to see what is happening.
In one, a well-designed AI tutor helped students learn more than twice as much as a lively active-learning classroom, in less time. In the other, students who used AI on their homework watched their exam scores fall even as their homework grades rose. The technology was the same. The outcome flipped on whether the student still did the mental work or handed it over.
Hold both results at once and, as we read it, the argument about whether AI is good or bad for learning dissolves. AI is an amplifier. Point it at effort and it multiplies learning. Point it at avoiding effort and it multiplies the avoidance.
04
2. The good case: a tutor that beats the classroom
We start with the good news, because it is genuinely good and easy to lose in the doom.
Gregory Kestin and colleagues at Harvard report a randomised controlled trial, published in Scientific Reports in June 2025, in a real undergraduate physics course, with 194 students. Every student experienced both conditions, an AI tutor and a well-run active-learning class, on different topics. On the test afterwards, students scored about 30 percent higher after the AI tutor and learned more than twice as much in less time.
The detail that matters is why. The tutor was not a chatbot told to help. It was engineered around decades of learning science: it made students attempt an answer before revealing anything, gave feedback in small steps, and kept them retrieving and explaining rather than passively reading. It manufactured what psychologists call desirable difficulty, the productive struggle that builds memory. The AI did not remove the effort. It aimed it.
We draw a lesson narrower than "AI tutors are good". What works is an AI built to make you think. What fails is an AI built to save you thinking. The same underlying model sits inside both the Harvard tutor and the homework shortcut in the next section, which is the whole point.
05
3. The crutch: the grades went up, the learning went down
Now the other side, at scale.
Economists at CEPR report tracking 26,811 Chinese secondary-school students over 30 months in 2026, as generative AI spread through their homework. The headline looked like a triumph. Homework scores rose about 18 percent and homework time fell by 30.
Then the exams came. Within six months, monthly exam scores had fallen around 20 percent. High-stakes entrance-exam scores dropped between 18 and 24 percent, and the full penalty only became visible after about two years. The losses were largest among the students you would least expect: the high achievers and the younger ones.
The mechanism is in the data. Roughly 80 percent of AI users showed the signature of outsourcing, very short homework time paired with high homework marks. The homework got done and the learning it was meant to produce did not happen. The tell is decisive: students who kept their homework time steady, and the effort with it, lost almost nothing. Same tool, same subjects, and the only variable that mattered was whether the mind stayed in the loop.
We will do the arithmetic in public, because the trade only becomes visible when the three numbers are put in one line. Homework time fell 30 percent. Homework scores rose 18 percent. Exam scores fell 20 percent within six months and entrance-exam scores between 18 and 24 percent over two years. On a nominal school workload of, say, ten hours of homework a week, the saving is three hours. What those three hours bought, at the top of the range, was a fifth to a quarter off the examination that determines which university a seventeen-year-old attends. Stated that way, almost nobody would take the trade. Stated as it was actually experienced, an evening that finished earlier and a mark that went up, almost everybody did.
That gap between how a decision feels and what it costs is the reason this is a design problem rather than a discipline problem. Nobody in that sample chose worse exam results. They chose a shorter evening, repeatedly, and the results were a consequence they could not see for two years.
Two things about this study deserve stating plainly, because it carries a great deal of weight in this dossier and in the wider debate. It is a working paper rather than a peer-reviewed publication. And it is observational: AI use was not randomly assigned, so students who chose to outsource may differ from those who did not in ways the controls cannot fully capture. The two-year lag and the dose-response pattern make the causal reading plausible. They do not make it certain.
06
4. Cognitive debt: the brain on ChatGPT
If the China study measured the outcome, an MIT Media Lab experiment watched the mechanism inside the skull.
Fifty-four people wrote essays wearing EEG caps, split into three groups: one used ChatGPT, one used a search engine, one used only their own head. The brain-only writers showed the strongest, most connected neural activity. The search users showed less. The ChatGPT users showed least.
The most striking result was not the brainwaves. It was the recall. More than eight in ten of the ChatGPT users could not quote a single sentence from the essay they had just finished. They had produced the text without encoding it. The researchers named the pattern cognitive debt.
This study is also the most over-quoted piece of evidence in the entire field, and section 10 is largely about that.
07
5. What the newer work refines
Our July rule was "effort is the variable". The work published since suggests we were close and slightly too coarse, and the refinement is practically useful.
A longitudinal analysis of student prompting strategies, published in 2026, finds that the split is not simply how much effort a student expends but what kind of cognitive operation they perform on the model's output. Students who actively evaluate and extend what the AI produces outperform those who accept it. Using the model as an information source, delegating the evaluative judgement to it, and staying in lower-order processing all associate with weaker performance. Two students can spend the same hour with the same tool and end up in different places depending on whether that hour was spent generating or judging.
That maps directly onto the finding in our what-stays-human dossier that the scarce human skill is vouching rather than producing, arrived at from an entirely different literature. When two lines of evidence converge on judging-over-generating from education research and from enterprise deployment data, we give it more weight than either alone.
A second strand replicates the China result in a controlled setting. A 2026 paper on mathematics problems reports the same shape in miniature: generative AI reduced study time and reduced the knowledge built, with faster completion and less learning in the same breath.
So I sorted the evidence base by size before going further, because the debate treats these studies as interchangeable and they are not. The China result rests on 26,811 students over 30 months. The Harvard tutor result rests on 194 students in a crossover design, where each student serves as their own control, which buys a lot of statistical power per head. The MIT neural result rests on 54 people, 18 per condition, over a single session. The clinician deskilling finding is one study at one site with a 6 percent effect measured 3 months in. The Wharton and Gerlich results are of the order of a few hundred participants each.

Put crudely: one study in this dossier has a sample large enough to detect a subtle effect in a messy real-world setting, one has a design tight enough to trust a large effect in a controlled one, and the rest are early indications. That ordering should govern how much weight each carries, and in the July edition it did not.
And a third strand is quietly the most significant, because it turns this from an argument into a measurement. Researchers have begun proposing instruments to quantify AI reliance directly, one of them scoring offloading by comparing a person's actual workflow against a counterfactual one without the tool. The July edition listed "longitudinal skill data" as something to watch. The more precise thing to watch is whether these instruments get adopted, because a school or an employer that can measure offloading can manage it, and one that cannot is guessing.
08
6. Never-skilling: the skill that never forms
There is a version of this worse than forgetting, because there was never anything to forget.
In July 2026 a Nature Medicine paper named it: never-skilling. It sits beside two older worries. Deskilling is an experienced clinician whose ability fades from disuse. Mis-skilling is a trainee who absorbs an AI's confident error as fact. Never-skilling is the quietest and most permanent of the three: a trainee lets the model generate the first diagnosis every time and so never builds the reasoning, because the AI did the cognitive work during precisely the window when the skill should have formed.
The authors are careful and so are we. Direct evidence from live medical training is thin, and the argument rests on established learning theory plus early signals from adjacent fields. Anthropic's own May 2026 research on how AI assistance affects the formation of coding skills is one of those adjacent signals, and it is notable mainly because a frontier lab chose to publish on the question at all.
The reason it matters beyond medicine is that the logic is general. Every profession has a formative window, the junior years when judgement is built by doing the boring, hard, first-pass work. Hand that work wholesale to a machine and you get output today and no expert in ten years. Our entry-level dossier traces the same fear through the labour market; this is its cognitive root.
09
7. Engagement is the variable
Put the studies side by side and one rule falls out, which section 5 has now sharpened.
When AI removes the cognitive operation, it builds debt. When AI adds desirable difficulty and leaves the judging to the person, it builds skill. The homework shortcut, the ghost-written essay and the auto-generated diagnosis all strip out the productive struggle, so the grades stay up while the understanding drains away. The Harvard tutor, the flashcard that forces retrieval and the assistant that asks you to explain your reasoning all keep the human doing the part that forms the skill, with the machine aiming the effort rather than absorbing it.
Which is why we think "should we let people use AI" is the wrong question. Everyone will use it. The real question is whether the AI in front of your students, your trainees or your team is designed to make them think or to save them thinking, and at the moment that is mostly being decided by accident, by whichever tool happens to be most convenient.
10
8. It does not stop at school
It would be comfortable to file this under education. The workplace evidence says otherwise.
Shaw and Nave at Wharton report in 2026 what they call cognitive surrender: professionals adopting AI output with minimal scrutiny, overriding both intuition and deliberation. The result was a sharp asymmetry. When the AI was right, the workers did better than they would have alone. When the AI was wrong, they did worse than people with no AI at all, because they had stopped checking.
The clinical version is the most concrete. One study reports that three months after AI assistance was introduced, doctors' ability to spot tumours without it had dropped by about 6 percent. The skill did not vanish; it eroded from disuse, exactly as the theory predicts. Gerlich at SBS Swiss Business School reports the heaviest reliance, and the lowest critical-thinking scores, among the youngest workers, the ones still forming the judgement they will lean on for the next forty years.
The classroom result and the office result are the same result.
11
9. The assessment problem
One casualty deserves its own section, because it breaks a tool every organisation relies on.
If AI can produce the essay, the report, the case write-up or the code review, then those artefacts no longer prove that a person can do the thing. The output stopped being evidence of the skill. Schools discovered this first, with the death of the take-home essay. Every company that evaluates people by what they hand in is about to discover it too.
Banning the tools never works and pretending nothing changed is worse. The answer we would build moves assessment back to where the thinking is visible: live problem-solving, explaining a decision out loud, defending an argument, doing the work in the room. That is more effort to run, and it is the only kind of assessment AI cannot quietly complete for the candidate.
This has stopped being theoretical. In July 2026 the University of Chicago Law School published an AI strategy built around three themes, the first of which is developing AI-resilient pedagogy and assessment, alongside elevating the human skills that distinguish excellent lawyers. Whatever one thinks of any particular institution's plan, it is the shape of the response we have been asking for, and it is now on paper somewhere that other institutions will read. Organisations that keep grading the artefact will keep promoting the tool instead of the person.
12
10. Where we could be wrong
Every other dossier in this series carries a section that undercuts its own argument. This one did not, and it should have, because the case against it is stronger than the July edition allowed.
The most-quoted study is the weakest one. The MIT EEG experiment has 54 participants, split three ways, which leaves 18 people per condition. It is a preprint, and it measures essay-writing under laboratory conditions over a short window. Mollick, who is not a sceptic about AI's risks, published a direct pushback in July 2025 under the title "Against Brain Damage", arguing that the study is being read far past what it can support and that the framing of AI as neurologically harmful is itself revealing.
I want to be specific about how this went wrong, because we participated in it. I went back through how the MIT result reached us and the trail is short: a headline finding, a striking statistic about recall, and a phrase, cognitive debt, that is memorable enough to travel without its sample size attached. I did not check the sample size before I quoted it, and neither did anyone the finding passed through on the way to us. At no point did we look at 18 people per arm and ask whether that supports a claim about how a technology affects human cognition. It does not. It supports a hypothesis worth testing at scale, which is what it was. We rated it medium-high in July and it is rated low-medium in the appendix here, and that change is a correction rather than new information: nothing about the study altered, only our reading of it. Section 4 is suggestive; section 3 is the evidence.
Offloading is not new, and it is not always bad. Writing offloaded memory. Calculators offloaded arithmetic. Satellite navigation offloaded route-finding, and we accept that most people traded a real skill for a real convenience and would not go back. The question we would ask is not whether a skill erodes but whether the eroded skill still matters, and for a good deal of what AI absorbs the answer will be no. A dossier that treats every instance of offloading as a loss is smuggling in a value judgement that should be argued rather than assumed.
The comparison group is usually flattering. Studies contrast AI-assisted learning against good instruction. Much real teaching is not good instruction, and for a student whose alternative is no help at all, a mediocre AI tutor may be a clear gain. The Harvard result also came from a purpose-built tutor with expert design behind it, which is not what most learners encounter, so it overstates what generic tools deliver.
And the causal chain has a weak link. The China study is observational, the never-skilling paper is a concept paper whose authors say the clinical evidence is limited, and the deskilling finding is a single result in one setting. The direction of the evidence is consistent, and that is why we still hold the argument. The individual pieces are less solid than the pattern they form, and anyone quoting one of them as settled is overreaching.
What survives all of that: the China study's dose-response relationship, where students who kept their effort kept their scores, is difficult to explain any other way, and it is the load-bearing finding in this dossier.
13
11. Design for capability, not convenience
Here we stop describing and start prescribing.
The entire market pressure of AI runs towards convenience: fewer clicks, less friction, the answer handed over. Learning runs the other way. It needs the right friction, at the right moment, on purpose. Those two forces pull against each other and, left alone, convenience wins every time, because convenience is what the product managers are measured on.
Designing for capability means we decide deliberately where a human must still do the work. That means using AI to remove the friction that teaches nothing, the formatting, the boilerplate, the search, and protecting the friction that builds judgement, the first attempt, the explanation, the decision under uncertainty. Section 5 makes this sharper than it was, because the operation to protect is the evaluating and extending rather than the typing.
In practice that is a redesign of how work and training are structured, not a policy memo about tool access. It is the same See, Understand, Adopt logic: see what the tool actually does to the skill, understand which friction is productive, then adopt a design that protects it.
Our one-line brief: use AI to delete the effort that teaches nothing, and protect the effort that teaches everything.
14
12. Your own people: reskill or de-skill
Every leader rolling AI into an organisation is running the China experiment on their own staff, whether they mean to or not.
Hand analysts a tool that writes the analysis and you get faster reports now, and in two years analysts who can no longer analyse. Hand them a tool that pressure-tests their thinking and you get the speed and sharper people at once. The difference will not show up in this quarter's output. It shows up the day the market shifts and you need judgement the tool cannot supply.
The part we find uncomfortable is that the de-skilling path looks better in the short run. It is cheaper, it is faster, and the metrics all improve, right up until the capability you quietly stopped building is the one you suddenly need. Note also that the China study's penalty took about two years to become fully visible, which is longer than most people stay in a role and considerably longer than most reporting cycles. The person who makes this decision will usually not be the person who pays for it.
So we would put a skills thesis inside any serious AI adoption plan: which capabilities in this organisation must get stronger rather than merely faster, and how the tools are set up to build them. That is a change-management question, and it is the one most rollouts never ask.
15
13. What we are watching
Whether offloading gets measured. Section 5's instruments are the development that would change this field from argument to management. An employer who can score reliance can act on it.
Longitudinal skill data. Most studies are still short. The two-year lag in the China result suggests the real costs arrive late, which means the current generation of reassuring short-run findings may simply be early.
Assessment redesign. Which schools and firms move from grading the artefact to observing the thinking. UChicago Law is one datapoint; the question is whether it becomes a pattern.
Productive-friction tools. The Harvard tutor is a template. Whether the market rewards AI built deliberately to make people retrieve and struggle, against AI built to be convenient, is a commercial question with an educational answer.
Rules for the formative window. Regulation on when AI may and may not be used during training and examinations. Medicine and law will likely move first.
16
14. Verification and sources
This dossier draws on live web research and a personal archive of more than 15,000 sources. The notes below flag confidence and the material caveats.
| Claim | Confidence | Note |
|---|---|---|
| Harvard AI tutor: ~30% higher test scores, more than twice the learning in less time, 194 students | High | Kestin et al., Scientific Reports, June 2025; crossover randomised trial in a physics course. Authors caution against generalising to higher-order synthesis. A purpose-built tutor, not a generic chatbot, which section 10 notes is a favourable comparison. |
| China: 26,811 students, homework scores +18% and homework time −30%, exam scores −20% within six months, ~80% showing the outsourcing signature, ~2-year lag to full effect | Medium-high | Stromberg, Lei & Wu, "The Generative AI Learning Penalty", CEPR DP21577, 2026. A working paper, not peer-reviewed, and observational rather than randomised. Large sample over 30 months and a clear dose-response pattern, which is why we lean on it; the causal reading remains an inference. |
| Entrance-exam scores fell 18-24%; losses largest among high achievers and younger students | Medium | Same study, subgroup effects. Subgroup results are less robust than headline results as a general rule. |
| MIT: weakest EEG connectivity among LLM writers; more than 80% could not quote their own essay | Low-medium | "Your Brain on ChatGPT", MIT Media Lab, arXiv 2506.08872, 2025. 54 participants, a preprint, short window, laboratory task. Downgraded from the July edition, which rated this medium-high. Section 10 explains why: it is the most over-quoted finding in this field, and Mollick's "Against Brain Damage" (July 2025) is a fair corrective. |
| Cognitive engagement, specifically evaluating and extending output rather than accepting it, separates gains from losses | Medium | Longitudinal analysis of student prompting strategies, 2026. Converges with independent findings in our what-stays-human dossier, which raises our confidence in the pattern without validating the individual study. |
| Generative AI reduced study time and knowledge built on mathematics problems | Medium | 2026 arXiv paper. A preprint; consistent with the China result in a controlled setting. |
| Instruments now proposed to score AI reliance by comparing actual against counterfactual workflows | Medium | 2026 arXiv paper. Proposed rather than validated or adopted; reported as a development to watch, not a tool to use yet. |
| "Never-skilling" as a distinct risk from deskilling and mis-skilling | High as a framework, low as evidence | Nature Medicine, July 2026. A concept paper. The authors state directly that clinical evidence is limited, and we repeat that rather than bury it. |
| Anthropic research on how AI assistance affects the formation of coding skills | Medium | Anthropic, May 2026. Research published by a company with an interest in the answer; cited as an adjacent signal, not as independent confirmation. |
| Wharton "cognitive surrender"; when the AI was wrong, assisted workers fell below the unassisted baseline | Medium-high | Shaw & Nave, Wharton, 2026. |
| Clinicians' unaided tumour detection dropped ~6% three months after AI support was introduced | Medium | A single deskilling study, context-specific, not replicated to our knowledge. |
| Heavier AI reliance correlates with lower critical thinking, youngest most affected | Low-medium | Gerlich, SBS Swiss Business School. Correlational with a self-report component; the direction of causation is not established. |
| UChicago Law published an AI strategy built on AI-resilient pedagogy and assessment | High | Published July 2026. An institutional plan, not an outcome; whether it works is unknown. |
| Offloading is longstanding and not always a loss (writing, arithmetic, navigation) | Framing | Our argument, offered in section 10 against our own case rather than as a finding. |
The framed figure in this dossier is taken from publicly posted material, credited to its original source in the caption. Charts labelled "BFF" are our own, drawn from the sources named beneath them.
Dossier as pdf
Download this dossier
The full dossier as a pdf, with every figure and the source list. Fill in your details and the download starts right away.