Unsealed filings in NYT v. OpenAI/Microsoft copyright suit reveal internal admissions calling AI training data scraping "theft" and an "existential threat" to publishers
Previously redacted material in the NYT v. OpenAI/Microsoft copyright lawsuit was unsealed, revealing internal company communications and documents: Microsoft's Brent Hecht called AI training data scraping "the largest theft of labor in human history" and "an astonishing theft of unprecedented proportions"; OpenAI's Nick Turley called publishers' situation an "existential threat"; Microsoft data showed Copilot caused up to 93% drop in click-through rates to NYT's domain; OpenAI's mid-training datasets contained 91,692+ copies of NYT/Daily News/CIR works; a Common Crawl-derived dataset had 2M+ nytimes.com documents; Project Mango dataset contained 160,903+ unique publisher works; internal messages showed OpenAI staff discussing a paywall-bypass "hack" approved by Sam Altman's co-founder Greg Brockman ("ah nice"); and researchers deliberately stripped copyright notices from training data.
Entities: Microsoft, OpenAI, The New York Times, New York Daily News, Center for Investigative Reporting, Brent Hecht
0 primary
What happened
Previously sealed filings in the NYT v. OpenAI/Microsoft copyright case were unsealed, revealing internal quotes and figures. Microsoft researcher Brent Hecht reportedly called AI training scraping "the largest theft of labor in human history"; OpenAI's Nick Turley called the threat to publishers "existential." Disclosed figures include a claimed 93% drop in click-through rates from Copilot to NYT's site, 91,692+ copies of NYT/Daily News/CIR content in OpenAI's mid-training data, and a Common Crawl-derived set with 2 million-plus nytimes.com documents. No primary source documents (the actual court exhibits) are linked or independently verified in this coverage.
Why it matters
This goes to the heart of the fair-use "substitution" argument central to nearly every AI copyright case: if internal communications show companies knew their products displaced demand for the content they trained on, that undermines a key defence. Publishers, other rights holders and regulators watching this case will treat it as a template for discovery strategy in parallel suits. Practically, though, nothing has legally changed yet: no ruling, no settlement, no damages awarded, so any business decisions based on this should wait for the court's actual findings.
What is noise
The quotes come from NYT's own advocacy brief, not from neutral court findings, and the underlying exhibits remain sealed, so context around lines like "ah nice" or "largest theft of labor" is genuinely missing and could read differently in full. Framing this as a admission of guilt overstates it: these are excerpts selected by an opposing party's lawyers to support summary judgment, not a verdict or settlement.
Watch next
- 01The unsealed exhibits themselves (not just NYT's characterisation of them) once fully public, to check whether quotes like Brockman's "ah nice" and Hecht's "largest theft of labor" line hold up in full context
- 02The presiding judge's rulings on OpenAI/Microsoft's fair-use summary judgment motion in NYT v. OpenAI, which will show whether this evidence actually moves the legal needle
- 03Whether other publisher plaintiffs (Daily News, CIR, and any new claimants) file amended complaints or new suits citing these specific documents and figures
Coverage
1 storyMore regulation signals
Full feed →- Cloudflare mandates AI companies to separate web crawlers for search and training1 Jul 202690
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Apple ships Gemini-powered Siri beta with iOS 27, excluding EU and China at launch15 Sept 202678