The defense isn't "we didn't use it" — it's "we barely got caught using it."

The Summary

The Signal

Microsoft's argument in its defense against the New York Times isn't that its AI doesn't reproduce copyrighted content. It's that Copilot rarely does it. That's not a denial. That's a risk assessment presented as exoneration. The framing matters because it shifts the conversation from "did you steal" to "how much stealing counts as stealing."

This isn't an isolated skirmish anymore. Seattle Times and Newsday are now suing OpenAI and Microsoft for the same alleged infringement, joining the NYT in what's becoming a coordinated legal offensive by publishers. Regional papers don't typically have the war chest or appetite for protracted litigation against Microsoft unless they smell blood in the water or see a class-action freight train they want a seat on.

"The case could redefine AI's role in content creation, impacting copyright laws and tech companies' responsibilities in using published material."

Here's what's actually at stake: every foundation model depends on ingesting massive corpora of text without asking permission first. The training paradigm is "scrape now, negotiate never." If courts decide that even rare verbatim reproduction crosses the line, or that training itself constitutes infringement, the entire LLM industry faces either a licensing crisis or a complete architectural rethink. Microsoft's "rarely" defense suggests they know exact reproduction is indefensible. They're betting judges will accept a probabilistic threshold for copying, the way search engines got safe harbor for indexing.

But search engines return links and snippets. LLMs return synthesized answers that can replace the original source entirely. The economic harm isn't theoretical for publishers:

  • Users who would have clicked through to read the full article now get the summary for free
  • Advertising revenue evaporates when the answer appears inline
  • Attribution becomes invisible when the model doesn't cite its sources

The tension between media outlets and tech companies isn't just about past infringement. It's about who controls the training data for the next generation of models. If publishers win, OpenAI and Microsoft either pay licensing fees that could run into billions, or they pivot to synthetic data and user-generated content, neither of which has the editorial rigor or factual grounding of professional journalism.

The Implication

Watch for settlement structures that look less like damage awards and more like ongoing revenue shares. Microsoft can afford to lose this case. It cannot afford the precedent that all past training runs were illegal and all future ones require pre-licensing every source. The real outcome will likely be a deal where publishers get a cut of AI revenue in exchange for retroactive permission and forward-looking access.

For anyone building agents or training models: the "move fast and apologize later" window is closing. The cost of apology is about to include a price list.

Sources

Crypto Briefing | Crypto Briefing