The AI industry's worst-kept secret just got a paper trail, and it's signed by the people building the models.

The Summary

The Signal

The documents don't reveal a smoking gun. They reveal the entire firearms cache. A Microsoft executive, writing internally, used language typically reserved for historical atrocities to describe the foundational practice of modern AI development. Then both companies kept doing it.

The unredacted filings show Microsoft and OpenAI built training datasets by systematically scraping content from behind the Times' paywall. Not public web content. Not Creative Commons material. Paid subscriber content that readers exchange money to access. The companies knew this. Internal communications show they understood they were taking work that cost money to produce and money to access, and using it without permission or payment.

"The largest theft of labor in human history" isn't activist rhetoric. It's what the people actually building these systems called it in private.

The timing matters. These conversations happened while both companies were publicly positioning themselves as responsible AI leaders. Microsoft published AI ethics principles. OpenAI's charter spoke of broadly distributed benefits. The gap between the public ethics framework and the private "take everything we can until someone stops us" strategy is now documented in court records.

Here's what the internal warnings said would happen:

  • Publishers would lose the ability to monetize their archives
  • The subscription model funding journalism would collapse
  • Training on this scale would create models that could reproduce paywalled work

All three predictions are playing out. The Times isn't suing because they're worried about the future. They're suing because ChatGPT can already reproduce their articles nearly verbatim, and every query that gets answered by the model instead of clicking through to Times.com is revenue that evaporated.

The legal argument is straightforward. Copyright law says you can't copy someone's work without permission and build a commercial product from it. The AI industry's defense has been that training is "transformative use" under fair use doctrine. But fair use gets harder to argue when your own executives are calling it theft in writing.

The Implication

The foundation of every frontier model is now legally vulnerable. Not speculative vulnerability. Active litigation with discovery producing documents that make the defense harder to mount. If the Times wins, every major AI lab faces the same liability for every paywalled source they trained on.

Watch for two paths forward. First, the retroactive licensing scramble. AI companies will race to cut deals with publishers, newswires, and archives to legitimize training data after the fact. Expect these deals to be expensive and exclusive. Second, the data sourcing shift. Future models will train on explicitly licensed content, synthetic data, and whatever public domain material hasn't been strip-mined yet. The free-for-all era of AI training is ending whether the lawsuits succeed or not, because no company wants their executives' "largest theft" emails read in open court.

For anyone building on these models: your foundation has cracks in it. Plan accordingly.

Sources

TechCrunch AI