H
Howardism
Plate IIAI Economics & LaborHOWARDISM

Experimental Learning Impact of Generative AI

PublishedJuly 16, 2026FiledConceptDomainAI Economics & LaborTagsGovernanceWorkforceEducationHuman CapitalHuman AI CollaborationEmpiricalReading11 minSourceAI-synthesised

Contractor & Reyes (arXiv 2607.08849): a randomized, proctored experiment with 211 undergraduates finds off-the-shelf AI access raises immediate test scores +0.27 SD, ~76% of which persists a week later on unaided tests, and lifts essay quality only after AI is removed — but the durable gains belong almost entirely to 'augmentation' users (AI as tutor/explainer) while 'automation' users' (AI-drafts-the-text) short-run gains vanish once AI is gone; the objective, measured-skill counterpart to the AEI self-report that learning both persists and can be hollow depending on use mode

Illustration for Experimental Learning Impact of Generative AI

Sources#

Summary#

Zara Contractor and Germán Reyes (Middlebury College, arXiv 2607.08849, July 2026) run the study most of the AI-and-learning corpus was missing: a randomized controlled experiment with an objective, proctored, unaided skill measure. 211 undergraduates learn an unfamiliar technical topic (blockchain, carbon capture, or CRISPR) and write an analytical essay in a 35-minute lab session, randomly assigned to AI-allowed or AI-forbidden; one week later all of them return and are tested and write again without any AI or resources. The headline: AI access raises immediate test scores by 0.27 SD, about 76% of that gain still shows up a week later unaided, and essay quality rises only after AI is removed. But the durable part is not uniform — it belongs almost entirely to students who used AI to explain concepts (augmentation), while students who used it to draft their text (automation) post large short-run essay gains that vanish completely once AI is gone. This is the near-controlled test of "outsource your thinking, not your understanding" and the objective-skill measure the automation–optimism link's self-report could not supply.

Evidence note. empirical and causal (randomization, F-test balance, double-lasso controls, ITT + TOT/2SLS) — the strongest design in this vault's learning cluster. But scope it honestly: elite liberal-arts undergraduates (mean GPA 3.68, SAT 1386, >80% already AI-adopters), a proctored 35-minute lab task, off-the-shelf ChatGPT (GPT-4o), one topic, one week. It measures learning per unit of time with time-on-task held fixed (AI changed how the hour was spent, not its length — plausibly because the lab offered few competing uses of time). The authors explicitly do not claim that real-world AI adoption raises learning overall: outside the lab students choose how long to study, and many use AI to save time. It informs, but does not settle, workplace skill-atrophy.

The design in one paragraph#

Between-subjects randomization at the lab level (eight time slots × two parallel labs). Session One: baseline 5-question test → 35-minute learning phase (read, search, draft a ~500-word essay; AI-allowed group gets a logged-in ChatGPT tab, AI-forbidden gets Google) → unaided 5-question post-test. Compliance enforced by proctors, ChatGPT logs, and interface screenshots. Session Two (~7 days later, everyone unaided): 10-question test + a 20-minute essay on the same topic, complementary prompt. Learning is measured two ways: knowledge tests (factual/conceptual recall) and essays (higher-order analysis), the latter graded blind by 311 master's/PhD Prolific graders plus an LLM grader, averaged, with objective linguistic features (length, readability, lexical diversity, textual similarity, and Pangram AI-detection) alongside.

First stage was strong: assignment lifted AI use by 67.3pp off a near-zero control; 68% of the treated used AI; 88% found it helpful. In the logs, the most common use was explaining concepts (57%), then drafting (31%), then summarizing (17%).

Durable test-score gains, concentrated in the middle — and skewed to the able#

  • Immediate: +6.7pp on a 56.3% control baseline = +0.27 SD (ITT; TOT +10.0pp / +0.40 SD). More than double the 0.10 SD median of Kraft (2020)'s 747 education RCTs, comparable to the 0.29 SD of structured human tutoring, ≈ a one-SD jump in GPA, ≈ $8,600/student of school spending.
  • Retention (one week, unaided): +5.1pp = +0.27 SD, ~76% of the immediate effect persists (the two are statistically indistinguishable). Durable, partially decaying knowledge.
  • Shape: the gain sits in the middle of the score distribution — the tails (0/1 and near-perfect scorers) barely move (Figure 4).
  • Who gains: larger in the upper GPA/SAT quartiles (bottom quartile ~0.05 SD; upper quartiles 0.19–0.40 SD) — so AI access may widen learning gaps, a genuine tension with the floor-raising story of software democratization.
  • Felt vs. measured: self-assessed knowledge shows no treatment effect (β=0.02). Students learned more but did not feel they had — the mirror image of the felt-vs-measured worry the AEI self-report leaves open.

Higher-order skills surface only after AI is removed#

In Session One, treated essays are longer, simpler, and 12.3pp more AI-flagged (a 96% jump), but overall quality rises only slightly and imprecisely (+0.17 SD, p=0.25) — because the essay is partly AI-written, blending output with learning. Notably, no homogenization: within-group essay similarity is flat, contra the convergence Brynjolfsson et al. (2025) found in workplace writing (open-ended prompts admit many valid answers).

The clean read comes from Session Two, written unaided a week later (Figure 6): the AI-detection effect collapses to 0.0pp and the stylistic differences fade — proving the Session One style shift was AI text entering essays, not durable change — while quality gains emerge: writing style & clarity +0.30 SD (p=0.016) and relevance to prompt +0.26 SD (p=0.041). The learning was real and it raised higher-order skill, not just fact recall — but it only became visible once the AI crutch was taken away.

Augmentation vs. automation: the controlled test of "outsource thinking, not understanding"#

The paper's most load-bearing result. Classifying treated users from their ChatGPT logs: 49% augmentation (AI works with the student — tutor, explainer), 32% automation (AI does the work — drafts the essay), 8% mixed, 11% off-topic. Validation that the split is real: automation users' essays are 53.6% AI-flagged vs 20.9% for augmentation users, and automation users spend 17% less time reading/searching.

Session 1 essay qualitySession 2 essay quality (unaided)Session 2 test (unaided)
Automation users+0.54 SD (p=0.029)+0.02 SD — vanishes+0.19 SD (p=0.33, n.s.)
Augmentation users+0.05 SD+0.22 SD (p=0.22)+0.29 SD (p=0.06)

Automation's Session One advantage is AI-produced output that does not survive its removal; augmentation's gains persist unaided, consistent with skill accumulation. This is a within-experiment demonstration that AI's short-run productivity boost and its long-run learning effect can point in opposite directions, decided by how the tool is used — the exact shape of Karpathy's thesis, and consistent with Strömberg et al. (2026) (AI raises homework, lowers exams, losses concentrated among outsourcers), Bastani et al. (2025) (base GPT-4: −0.19 SD later), and Shen & Tamkin (2026) (engineers who learned a library via AI scored worse unassisted). The meta-analysis (Figure 5, grand mean 0.18 SD across 22 estimates) sorts the whole literature along this axis: losses cluster where AI could do the practice for you, gains where AI plays coach.

Mechanisms#

  • Time reallocated, not saved. No effect on total learning time (unlike the large workplace time-savings of Noy & Zhang 2023), but the mix shifts −5.3pp writing / +4.4pp reading-and-searching — effort moves from producing text to absorbing it (task execution → stewardship, à la Copilot users spending half their time verifying).
  • Enjoyment up 13% (+0.66 pts, p=0.031): AI made learning more engaging, not mechanical.
  • Cheating up but not the driver: rule violations rose +12.6pp combined (p=0.005), yet a back-of-envelope bound attributes at most ~2.2pp (about a third) of the test-score gain to cheating — the effect is mostly real learning.

Beliefs: experience corrects the magnitude, not the direction#

Both groups expect AI to help, but only the treated gauge how much correctly. Control students predict a +25.2pp own gain — about 5× the actual 5.1pp; treated students, who used it, predict +3.8pp, close to the estimate. Perceived gains track actual gains across subgroups. And students already hold the right model: 69% name both a help and a harm channel, 54% say the effect "depends on how AI is used," and their two most-named mechanisms (AI explains/tutors 40%; AI shortcuts the work 42%) map exactly onto augmentation vs automation. Students grasp the mechanism from the outset; firsthand use only calibrates the magnitude — a belief-updating result that rhymes with the professional-misjudgment literature (Becker et al. 2025 / METR: developers predicted +24% speedup, measured −19%).

What it settles — and what it doesn't#

It settles, for this population and task, that off-the-shelf AI can produce durable, objectively measured learning gains, dissolving the "AI must erode learning" prior — but conditions the result on use mode. It does not settle: whether the gains hold outside a proctored lab where time is endogenous; whether they hold for non-elite learners (they skew to the able, suggesting not uniformly); or whether the same augmentation/automation split governs workplace deskilling, where the analogue evidence is still mixed (Brynjolfsson's support agents learn; Budzyń's endoscopists deskill).

Connections#

  • The Automation–Optimism Linkthe primary complement. The AEI survey found heavy delegators self-report no learning loss; this experiment supplies the objective, randomized measure that survey could not — and it both agrees (augmentation users' learning persists unaided) and exposes the divergence the survey can't see (automation users' gains are hollow, and "automation share" pools both types)
  • Outsource Your Thinking, Not Your Understanding — the near-controlled test of the thesis: augmentation (AI helps you understand) builds durable skill; automation (AI does the thinking) leaves nothing once removed — Karpathy's principle, randomized
  • AI Brain Fry — both put an objective, measured number on AI's cognitive effect (there, oversight fatigue → +11%/+39% errors; here, learning → +0.27 SD, or hollow gains for automators) against the softer self-report signal; the deskilling half of this paper is that page's mechanism in a learning task
  • Returns to Expertise in Agentic Coding — the heterogeneity rhymes: gains skew to higher-ability students, and augmentation ≈ using AI to deepen the understanding that Anthropic's study finds is what amplifies an agent; both say the benefit accrues to whoever brings (or builds) understanding
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the belief half: like reported/anticipated exposure, students' perceptions of AI's effect are directionally right but miscalibrated in magnitude until firsthand use corrects them
  • Printing Press Software Democratization — the tension: democratization predicts AI raises the floor, but here gains concentrate in the upper ability quartiles, hinting AI may widen learning gaps rather than close them
  • Configurable Human Participation — the system-side mirror, published the same week: HAS-Bench's agency scale runs on the same augmentation-vs-automation axis (A1–A2 automation-oriented vs A3–A5 augmentation-oriented), and both land the same shape of result — the value of AI–human collaboration is decided by its mode and timing, not its amount (there: more agency at A4 breaks tasks A3 solved; here: automation-mode use yields hollow gains that vanish with the tool)

Open Questions#

  • Time-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the augmentation dividend survive once students can spend the hour AI frees on something else entirely?
  • Gains skew to the able (upper GPA/SAT quartiles). Is the widening-gaps signal a durable property of unrestricted AI, or an artifact of a high-ceiling elite sample where the bottom quartile has little room to move?
  • The augmentation/automation choice is endogenous to incentives (grade inflation and signaling-motivated students push toward automation). Can incentive or interface design shift the mix toward augmentation at scale — and would that reverse the deskilling half?
  • Does the same use-mode split govern workplace skill accumulation (the open question The Automation–Optimism Link and AI Brain Fry leave for workers), or is a proctored one-week academic task too unlike on-the-job learning to transfer?

Sources#

  • Experimental Evidence on the Learning Impact of Generative AI — Contractor & Reyes, Experimental Evidence on the Learning Impact of Generative AI (arXiv 2607.08849, 2026-07-09), empirical. §4 immediate effects, §5 retention + augmentation/automation heterogeneity, §6 beliefs, §7 conclusion; Figures 4–8, Tables 4–8.
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 9
Related articles
  • The Automation–Optimism Link

    AEI Cadences survey finding: people who use Claude in more automated ways are MORE optimistic across all six job-qualit…

  • Organizational Complements to AI

    The general-purpose-technology argument that AI's productivity gains depend on complementary workflow/skill/org-design…

  • Returns to Expertise in Agentic Coding

    Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…

  • Unknowns as the Agentic Bottleneck

    Thariq Shihipar's map-vs-territory thesis: the gap between what you told the agent and what the work actually requires…

  • AI Brain Fry

    Kropp et al. 2026/03: mental fatigue from excessive AI oversight increases minor errors +11%, major errors +39%; cognit…