Signum
Feed
Useful signal17 Sept 2026medium confidence

Unsealed filings in NYT v. OpenAI/Microsoft copyright suit reveal internal admissions calling AI training data scraping "theft" and an "existential threat" to publishers

Previously redacted material in the NYT v. OpenAI/Microsoft copyright lawsuit was unsealed, revealing internal company communications and documents: Microsoft's Brent Hecht called AI training data scraping "the largest theft of labor in human history" and "an astonishing theft of unprecedented proportions"; OpenAI's Nick Turley called publishers' situation an "existential threat"; Microsoft data showed Copilot caused up to 93% drop in click-through rates to NYT's domain; OpenAI's mid-training datasets contained 91,692+ copies of NYT/Daily News/CIR works; a Common Crawl-derived dataset had 2M+ nytimes.com documents; Project Mango dataset contained 160,903+ unique publisher works; internal messages showed OpenAI staff discussing a paywall-bypass "hack" approved by Sam Altman's co-founder Greg Brockman ("ah nice"); and researchers deliberately stripped copyright notices from training data.

GovernanceEconomicsLabourAccess

Entities: Microsoft, OpenAI, The New York Times, New York Daily News, Center for Investigative Reporting, Brent Hecht

74Useful signal
1 source
0 primary
Was this useful?
01

What happened

Previously sealed filings in the NYT v. OpenAI/Microsoft copyright case were unsealed, revealing internal quotes and figures. Microsoft researcher Brent Hecht reportedly called AI training scraping "the largest theft of labor in human history"; OpenAI's Nick Turley called the threat to publishers "existential." Disclosed figures include a claimed 93% drop in click-through rates from Copilot to NYT's site, 91,692+ copies of NYT/Daily News/CIR content in OpenAI's mid-training data, and a Common Crawl-derived set with 2 million-plus nytimes.com documents. No primary source documents (the actual court exhibits) are linked or independently verified in this coverage.

02

Why it matters

This goes to the heart of the fair-use "substitution" argument central to nearly every AI copyright case: if internal communications show companies knew their products displaced demand for the content they trained on, that undermines a key defence. Publishers, other rights holders and regulators watching this case will treat it as a template for discovery strategy in parallel suits. Practically, though, nothing has legally changed yet: no ruling, no settlement, no damages awarded, so any business decisions based on this should wait for the court's actual findings.

03

What is noise

The quotes come from NYT's own advocacy brief, not from neutral court findings, and the underlying exhibits remain sealed, so context around lines like "ah nice" or "largest theft of labor" is genuinely missing and could read differently in full. Framing this as a admission of guilt overstates it: these are excerpts selected by an opposing party's lawyers to support summary judgment, not a verdict or settlement.

04

Watch next

  1. 01The unsealed exhibits themselves (not just NYT's characterisation of them) once fully public, to check whether quotes like Brockman's "ah nice" and Hecht's "largest theft of labor" line hold up in full context
  2. 02The presiding judge's rulings on OpenAI/Microsoft's fair-use summary judgment motion in NYT v. OpenAI, which will show whether this evidence actually moves the legal needle
  3. 03Whether other publisher plaintiffs (Daily News, CIR, and any new claimants) file amended complaints or new suits citing these specific documents and figures

Coverage

1 story

More regulation signals

Full feed →