SDSignal Desk

Microsoft says virtually nobody was grabbing NYT articles through its chatbot

Sep 4, 2026, 9:05 AM · The Verge

Image: The Verge

Microsoft is trying to turn discovery volume into a fair-use argument: millions of Copilot chats, almost no long matches. Publishers still call that commercial substitution, not a rounding error.

Why it matters

The copyright fight against Microsoft and OpenAI is moving from rhetoric into spreadsheet warfare. Microsoft says it turned over 8.2 million Copilot chat logs—selected because they hit keywords tied to news publishers' sites—and that the resulting expert work shows vanishingly little regurgitation of the works at issue.

That claim lands as Microsoft asks for summary judgment. If a judge buys the "transformative training, rare output" frame, early dismissal becomes plausible. If not, the consolidated publishers-and-authors case keeps grinding toward trial, with substitution and market harm still on the table.

The Signal Desk read

Microsoft's filing is a classic defense move: flood the record with volume, then argue the plaintiff's nightmare scenario almost never happens. Fewer than 1 percent of those 8.2 million logs, it says, shared at least 16 words with grounding news content—59,545 chats. An Authors Guild–side expert, per Microsoft, found only 24 responses with at least 30 matching words across the same pile, and matches in only 10 of 212 books evaluated. For the Center for Investigative Reporting, Microsoft cites an expert finding 51 instances of "substantial overlap."

Those numbers are designed to sound like proof that Copilot is not a substitute newspaper. They are also carefully scoped. The logs were the ones *most* likely to touch publisher material, Microsoft says—which makes the low hit rate rhetorically powerful, and also means the sample is not a random window into all Copilot use. Sixteen-word and thirty-word thresholds are blunt instruments; publishers care about whether a chat replaces a subscription session, not whether a string-match algorithm lights up.

The New York Times' reply is blunt and unsurprising: discovery supposedly shows theft into commercial products that substitute for journalism. That is the real clash of theories. Microsoft wants the court to treat occasional overlap as noise around a transformative purpose. The Times wants the court to treat training-plus-product as a business that free-rides on reporting and then competes with it. Frequency of verbatim dumps is relevant to one story and almost beside the point for the other.

The likelier read is that summary judgment is a stretch if the case turns on market substitution rather than copy-paste counts. Judges can accept that LLMs rarely spit out full articles and still find a genuine dispute about whether the systems are built on copyrighted works in a way the statute does not bless. Microsoft's numbers help on regurgitation; they do less work on the training-data theory that has always been the publishers' core grievance.

One more pressure point sits outside the expert tables: the Trump administration filed a statement of interest supporting OpenAI in the Times matter this week. Political wind does not decide fair use, but it does change how hard parties push settlement versus principle. Expect Microsoft to keep selling "virtually nobody," and expect publishers to keep answering that the product is the problem even when the paste is rare.

Context

News publishers and book authors sued Microsoft and OpenAI over training and product use of their works; the claims were consolidated before one judge over plaintiff objections. Microsoft's Friday filing is part of its push for an early end via summary judgment rather than a full trial on fair use and substitution.

Who feels it

Publishers and authors
They need to keep the case about training and market harm, not about how often Copilot dumps 30 contiguous words. If regurgitation becomes the whole fight, Microsoft's spreadsheet wins.
Microsoft and OpenAI
Low match rates are the best fact pattern they have for summary judgment. The risk is a judge who treats those rates as interesting and still finds a triable fair-use dispute.
Other AI companies watching the docket
A ruling that rare output overlap proves transformative purpose would travel far beyond Copilot. A denial that keeps training itself in play would do the opposite.

What to watch

  1. Whether the court grants Microsoft's summary-judgment bid or keeps the consolidated case on a path to further proceedings.
  2. How the Times and other publishers answer the 8.2 million-log analysis with their own substitution and market-harm evidence.
  3. Whether the administration's statement of interest reshapes settlement posture or remains a footnote to the fair-use briefing.

Read the original

Continue at the source.

The Verge

Companies: Microsoft