Will it run?
Policy

Microsoft says Copilot almost never reproduces New York Times articles

By Nadia Ksiazek Clawpit staff

Microsoft filed a court analysis of 8.2 million Copilot conversations, pre-selected because they contained news-related keywords — the sample the company considered most “dangerous.” Of those, 59,545 conversations, or less than one percent, contained at least 16 consecutive words identical to news content used in training. An expert for the Center for Investigative Reporting identified 51 instances of “substantial overlap” with CIR material. In the authors’ suit, only 24 responses had 30 or more matching words. Of 212 books examined, just 10 showed any match.

The Times rejects the conclusions outright. Lead counsel Ian Crosby said the documents and testimony uncovered in discovery lead to a single conclusion: Microsoft and OpenAI stole from the Times to build commercial products that compete with its journalism, threaten its business model and undermine the industry as a whole. The Center for Investigative Reporting and the Authors Guild did not respond to requests for comment.

Microsoft argues the data bolsters its fair-use defense: using copyrighted material to train language models should qualify as fair use. The systems rely on protected works but serve fundamentally different purposes from the originals. The fact that they occasionally reproduce text passages, Microsoft contends, “barely undermines the transformative purpose of LLM training.”

The motion was filed in the consolidated litigation brought by publishers and authors alleging the companies built products on their works and now compete with them directly, in part by reproducing protected content. Microsoft is asking the judge for summary judgment to end the case early. This week the Trump administration filed a statement of interest in the Times suit, backing OpenAI’s position. If the judge denies the motion, the trial proceeds.