Signum
Feed
Useful signal4 Sept 2026high confidence

Microsoft files legal analysis in NYT copyright suit claiming Copilot rarely reproduces substantial copyrighted text

Microsoft submitted a legal filing in the consolidated NYT/publishers/authors copyright lawsuit, presenting an analysis of 8.2 million Copilot chat logs showing fewer than 1% (59,545) contained at least 16 words matching news content, 51 instances of 'substantial overlap' with CIR work, and only 24 of 8.2 million author-suit conversations had 30+ matching words (10 of 212 books had any matches). The filing argues for summary judgment to end the case early.

GovernanceEconomicsAccess

Entities: Microsoft, OpenAI, The New York Times, Copilot, Center for Investigative Reporting, Authors Guild

67Useful signal
1 source
0 primary
Was this useful?
01

What happened

Microsoft submitted a legal filing in the consolidated NYT/publishers/authors copyright lawsuit against Microsoft and OpenAI, arguing for summary judgment. The filing presents Microsoft's own analysis of 8.2 million Copilot chat logs: fewer than 1% (59,545) contained at least 16 consecutive matching words from news content, 51 instances showed "substantial overlap" with Center for Investigative Reporting material, and only 24 of 8.2 million conversations in the authors' suit had 30+ matching words (affecting just 10 of 212 books checked). Nothing has been decided yet; this is Microsoft's argument for dismissal, not a court ruling.

02

Why it matters

This case will help set the legal boundaries for how much copyrighted material AI companies can be liable for when chatbots reproduce it, which affects training-data licensing costs and fair-use defences across the industry, not just Microsoft and OpenAI. If the summary judgment succeeds, it could weaken publishers' leverage in ongoing and future AI licensing negotiations. If it fails, it signals that even low reproduction rates may not shield companies from liability, raising the bar for content-filtering and licensing across the sector.

03

What is noise

The framing that "virtually nobody" was extracting NYT content uses a self-selected 16-word threshold chosen by the defendant, not an industry or judicial standard, so the low percentages should not be read as a settled fact about infringement. NYT's counsel disputes the conclusion, and discovery statistics from an interested party in active litigation are not the same as a judicial finding; treat the numbers as one side's argument, not a verdict.

04

Watch next

  1. 01The judge's ruling on Microsoft's summary judgment motion, expected in coming months, which would determine whether the case proceeds to trial
  2. 02NYT and Authors Guild's counter-filing or rebuttal disputing Microsoft's methodology and word-count thresholds
  3. 03Any parallel rulings or settlements in similar AI copyright suits (e.g. against OpenAI directly, or other publishers) that could signal how courts are treating chat-log discovery evidence

Coverage

1 story

More regulation signals

Full feed →