OECD PISA study finds AI-using students generally score lower on tests, except with moderate, critically-assessed use
The OECD's 2026 PISA report (based on 2025 data from over 760,000 15-year-olds in 91 countries) found that, after adjusting for socioeconomic status, students who never use AI generally outperform AI users in science; task type, frequency of use, and whether students were trained to critically assess AI output all significantly affect whether AI use correlates with better or worse test scores. Weekly AI use for general learning combined with critical-assessment training was associated with outperforming non-users.
Entities: OECD, PISA (Programme for International Student Assessment), Andreas Schleicher
0 primary
What happened
The OECD's 2026 PISA report, drawing on 2025 data from over 760,000 15-year-olds across 91 countries, found that students who never use AI generally outperform AI users in science, after adjusting for socioeconomic status. The relationship is not uniform: task type, frequency of use, and whether students were trained to critically assess AI output all changed the direction of the effect. Weekly AI use for general learning combined with critical-assessment training was linked to outperforming non-users, while unstructured or frequent use was linked to worse scores.
Why it matters
This is the largest dataset yet on AI use and learning outcomes, giving education ministries and school leaders a reference point for AI-in-classroom policy debates that have so far run on anecdote and vendor claims. The core implication for schools, edtech vendors and policymakers is that access to AI tools alone does not help students, and may hurt them, unless paired with explicit training in critical evaluation of AI output. That is a concrete argument for curriculum investment (AI literacy training) over device or license rollouts, which affects procurement and policy decisions in the next budget cycle.
What is noise
The Verge's headline ("students who use AI generally score worse") flattens a conditional, subgroup-dependent finding into a blanket claim, and Schleicher's "productive cognitive struggle" framing asserts a causal mechanism the correlational design cannot establish. No effect sizes or confidence intervals are reported here, and reverse causality is unaddressed: weaker students may turn to AI more, or better-supported students may be the ones getting critical-assessment training in the first place. The primary OECD report itself has not been linked or reviewed, so this briefing is one step removed from the underlying methodology and data tables.
Watch next
- 01Locate and review the primary OECD PISA 2026 report for effect sizes, confidence intervals, and how 'critical-assessment training' was measured, not just its existence
- 02Watch for independent replication or critique from education researchers (e.g. in journals or via groups like EEF, NBER) addressing reverse causality and confounding in the training subgroup
- 03Track whether education ministries (UK DfE, US states, EU members) cite this report in AI-in-schools policy or curriculum funding decisions over the next 6-12 months