The content licensing market for AI is booming, but publishers are discovering they're price-takers in a buyer's market they didn't see coming.

The Summary

The Signal

The AI training data market isn't functioning like a market at all. It's functioning like a hostage negotiation where one side already ate the hostages.

Internal documents from the New York Times lawsuit show Microsoft director Brent Hecht warning in early 2023 that mass content scraping would be seen as "an astonishing theft" and "the largest theft of labor in human history." OpenAI President Greg Brockman responded to a researcher bypassing the Times paywall with "ah nice." These weren't rogue engineers — these were executives documenting their awareness that what they were building depended on taking content without permission, compensation, or even conversation.

The lawsuit reveals OpenAI's head of ChatGPT calling these tools an "existential threat" to publishers. That's not speculation from anxious newsrooms. That's the builder's own assessment.

"Content has value, even if it's scraped 'for free' from the internet. But everybody agrees it's not zero."

Now publishers have what they supposedly wanted: AI companies willing to license content. The Times signed with OpenAI. Axel Springer made deals. News Corp got a reported $250 million from OpenAI over five years. But here's what changed between the scraping era and the licensing era: nothing structural.

The AI labs built their models on scraped content first. That means:

  • They already have most of what they need in their training sets
  • New licensing deals are marginal improvements, not existential dependencies
  • Publishers negotiate from a position of "pay us or we'll sue" rather than "you need us to build this"

This is price discovery in reverse. Usually markets find equilibrium through supply and demand. Here, demand was satisfied illegally, then legitimized retroactively through settlements and licensing deals that reflect buyer leverage, not content value.

The result is a market where publishers can't act collectively (antitrust concerns), can't withold supply (already scraped), and can't price based on replacement cost (models already trained). They can only accept whatever deal keeps them in the AI training pipeline for future model iterations.

The Implication

Watch for two-tier content pricing to emerge: premium deals for publishers with legal firepower (Times, Murdoch properties, big magazine groups) and take-it-or-leave-it terms for everyone else. The winners will be publishers who can prove their content is uniquely valuable for specific model capabilities — not just "more text."

If you're a content creator or publisher, the window for meaningful negotiating leverage is closing. The first generation of frontier models is already trained. The question isn't whether your content has value. It's whether you can prove it has *marginal* value when the buyer already took it once and might just wait you out for the copyright lawsuits to settle.

Sources

Fast Company Tech