The discovery logs just became the defense.

The Summary

  • Microsoft turned over 8.2 million Copilot chat logs in the NYT copyright lawsuit — logs hand-picked because they contained keywords most likely to trigger reproductions of publisher content
  • Out of those 8.2 million logs, only 59,545 chats surfaced any output resembling publisher material, and even fewer reproduced "substantive chunks"
  • Microsoft's argument: even when users are explicitly searching for news content, the models overwhelmingly paraphrase, summarize, or cite rather than copy

The Signal

This is the first hard data we've seen on what actually happens when millions of people use generative AI to interact with copyrighted material. Microsoft's filing claims that out of 8.2 million chat logs selected specifically because they mentioned NYT or other plaintiff publishers, fewer than 60,000 produced any output that even resembled the original work. That's 0.7 percent. And "resembled" doesn't mean "reproduced." Microsoft says even fewer gave users substantive chunks that could substitute for reading the actual article.

The logs weren't random. Microsoft chose them because they hit on keywords most likely to surface plaintiff content. They tilted the sample toward finding violations. If this is the worst-case set and it still shows 99.3 percent of chats doing something other than straight copying, that's a problem for the publishers' theory of harm.

"Even when optimized to find infringement, the logs show users aren't using Copilot as a piracy tool."

The real question is what those 59,545 chats actually contained. Did they spit out a NYT lede verbatim? A full paragraph? A sentence fragment that happened to match? Microsoft's framing suggests most were paraphrases or summaries, not substitutes. That distinction matters. If I ask Copilot "What did the NYT say about the Fed's rate decision?" and it gives me a three-sentence synopsis, that's closer to what a human research assistant does than what Napster did.

Publishers argue the models were trained on their archives without permission, so every output is tainted. Microsoft is arguing the outputs themselves prove no mass-scale substitution is happening. Both can be true. The training might be legally dubious. The usage might not be causing the harm plaintiffs claim.

The filing also reveals how discovery works in AI litigation. Plaintiffs get access to millions of real-world logs. That's a gold mine for understanding what users actually do with these tools, but it's also a risk. If the data doesn't support your theory, you've just funded the defense's best exhibit. Microsoft clearly believes the logs help them more than they help the NYT.

The Implication

If this data holds up under scrutiny, it shifts the fight. The publishers can't credibly claim Copilot is cannibalizing their traffic if less than one percent of targeted searches produce anything close to their original text. They'll need to pivot to a pure training argument, which is harder to win and offers less in damages. Watch for the plaintiffs to argue that even 59,545 instances of reproduction is 59,545 too many, or that Microsoft cherry-picked the logs. The discovery battle is now the case.

Sources

The Verge AI