Microsoft files legal analysis in NYT copyright suit claiming Copilot rarely reproduces substantial copyrighted text
Microsoft submitted a legal filing in the consolidated NYT/publishers/authors copyright lawsuit, presenting an analysis of 8.2 million Copilot chat logs showing fewer than 1% (59,545) contained at least 16 words matching news content, 51 instances of 'substantial overlap' with CIR work, and only 24 of 8.2 million author-suit conversations had 30+ matching words (10 of 212 books had any matches). The filing argues for summary judgment to end the case early.
Entities: Microsoft, OpenAI, The New York Times, Copilot, Center for Investigative Reporting, Authors Guild
0 primary
What happened
Microsoft submitted a legal filing in the consolidated NYT/publishers/authors copyright lawsuit against Microsoft and OpenAI, arguing for summary judgment. The filing presents Microsoft's own analysis of 8.2 million Copilot chat logs: fewer than 1% (59,545) contained at least 16 consecutive matching words from news content, 51 instances showed "substantial overlap" with Center for Investigative Reporting material, and only 24 of 8.2 million conversations in the authors' suit had 30+ matching words (affecting just 10 of 212 books checked). Nothing has been decided yet; this is Microsoft's argument for dismissal, not a court ruling.
Why it matters
This case will help set the legal boundaries for how much copyrighted material AI companies can be liable for when chatbots reproduce it, which affects training-data licensing costs and fair-use defences across the industry, not just Microsoft and OpenAI. If the summary judgment succeeds, it could weaken publishers' leverage in ongoing and future AI licensing negotiations. If it fails, it signals that even low reproduction rates may not shield companies from liability, raising the bar for content-filtering and licensing across the sector.
What is noise
The framing that "virtually nobody" was extracting NYT content uses a self-selected 16-word threshold chosen by the defendant, not an industry or judicial standard, so the low percentages should not be read as a settled fact about infringement. NYT's counsel disputes the conclusion, and discovery statistics from an interested party in active litigation are not the same as a judicial finding; treat the numbers as one side's argument, not a verdict.
Watch next
- 01The judge's ruling on Microsoft's summary judgment motion, expected in coming months, which would determine whether the case proceeds to trial
- 02NYT and Authors Guild's counter-filing or rebuttal disputing Microsoft's methodology and word-count thresholds
- 03Any parallel rulings or settlements in similar AI copyright suits (e.g. against OpenAI directly, or other publishers) that could signal how courts are treating chat-log discovery evidence
Coverage
1 storyMore regulation signals
Full feed →- Cloudflare mandates AI companies to separate web crawlers for search and training1 Jul 202690
- FERC mandates fast lane for data center interconnections to the grid18 Jun 202682
- Police officer investigated for using AI to create evidence in multiple cases13 Jun 202682
- OpenAI submits draft S-1 to the SEC8 Jun 202682