When the executive branch picks sides in a copyright war, it's not about fair use — it's about who gets to own the future of information.

The Summary

The Signal

The New York Times filed this lawsuit in December 2023, claiming OpenAI and Microsoft unlawfully used its journalism to train ChatGPT and seeking billions in damages. Now the federal government has entered the chat, and its position is clear: training AI on copyrighted text is fair use. That's not a small thing. This is the executive branch formally backing the legal theory that underpins the entire foundation model industry.

The DOJ's brief argues the Times is trying to "narrow fair-use doctrine" in ways that would exclude AI training entirely. Translation: the government sees AI model training as fundamentally transformative work, not just fancy copying. If that view holds, every publisher's leverage evaporates. You can't negotiate licensing deals if the law says the models don't need your permission in the first place.

"The DOJ is arguing that AI training is transformative enough to qualify as fair use — which guts the entire business model of content licensing to AI labs."

The timing matters. OpenAI is in the middle of multiple copyright battles, including one with authors and another with visual artists. Every case tests the same question: does scraping the internet to train a model count as fair use, or is it mass copyright infringement at scale. Federal backing tips the scale. Judges pay attention when the government files statements of interest, especially on unsettled law.

This isn't about Trump or Biden or any specific administration. It's about regulatory capture by the companies building the rails. OpenAI, Anthropic, Google, Meta — they all need the same legal framework to survive. They need training data to be free, or at least legally defensible. A ruling against them doesn't just cost money. It breaks the model.

Key implications for the agent economy:

  • If training is fair use, foundation models stay cheap to build and improve
  • If it's not, model development gets gated by licensing costs and legal risk
  • Smaller labs without legal teams or deep pockets get priced out entirely

Publishers are watching this case closely because they're trying to build their own leverage. Some have cut deals with OpenAI and others for training data access. But those deals only matter if the law requires them. If fair use wins, the deals were unnecessary. If infringement wins, everyone who didn't get paid has a case. The government just telegraphed which side it thinks should win.

The Implication

Watch what publishers do next. If they believe fair use will carry the day, expect them to shift strategy from litigation to differentiation — building products and services AI can't replicate by ingesting text. If they think they can still win in court, expect more lawsuits and more pressure on Congress to clarify copyright law for AI training.

For anyone building agents or tools on foundation models, this is good news in the short term. Training data stays accessible. Model costs stay manageable. But the long-term play is messier. If the legal framework stabilizes around "training is fair use," content creators lose a revenue stream. That changes what gets made, who makes it, and how much of it exists. Foundation models get trained on a corpus that stops growing, or grows slower, or skews toward whoever can still afford to create without getting paid for training use.

The real question isn't whether OpenAI wins this case. It's whether the legal system locks in a framework that treats information as free for the taking, or builds in mechanisms for creators to capture value when their work trains the machines. The government just showed its hand.

Sources

The Verge AI