The internet ran dry, so now they're strip-mining bankruptcy courts.

The Summary

The Signal

Google beat out Mercor, an AI training startup, by $2.5 million for Spirit's internal records. Not customer names or payment info. The operational guts: scheduling systems, maintenance logs, customer complaint workflows, pricing algorithms, crew management software. A Google spokesperson confirmed the data "can be helpful in improving our products and AI models" and clarified they're buying "internal data and custom software" but explicitly not customer or credit card information.

This is the new frontier. The public internet has been scraped clean. Every Wikipedia article, Reddit thread, and blog post is already in someone's training set. Now the valuable stuff is locked in corporate databases that were never meant to be public. As Business Insider notes, corporate data is "especially valuable because it can help train bots to handle customer complaints more effectively or debug websites."

"Corporate data is especially valuable because it can help train bots to handle customer complaints more effectively or debug websites."

Spirit Airlines ran one of the most operationally complex budget airline operations in the US before filing Chapter 11. That means:

  • Millions of customer service interactions across chat, phone, and email
  • Dynamic pricing models optimized for razor-thin margins
  • Crew scheduling algorithms that juggled FAA regulations with cost efficiency
  • Maintenance tracking systems tied to fleet operations

All of that is behavioral data about how a real business operated under real constraints. It's not synthetic. It's not cleaned up for public consumption. It's the messy, real-world decision-making that AI companies need to train agents that can actually do work, not just sound convincing in a chatbot demo.

Mercor, which runs a program paying midsize companies $100K for their data, got outbid because Google has deeper pockets and a clearer strategic need. Google isn't just training a chatbot. They're building Gemini models that need to understand how businesses actually run — how real companies handle logistics, customer escalations, inventory management, pricing decisions. You can't learn that from scraping the public web.

The Implication

Bankruptcy courts are now data auctions. Every company that goes under with intact databases is suddenly a training set for sale. Expect more of this. Startups will start bidding on failed retail chains, logistics companies, healthcare providers. The corporate graveyard is the new gold mine.

If you're running a company with proprietary operational data, you now have a sellable asset even in liquidation. If you're building AI agents meant to operate in specific verticals, your competitors are already buying up the best training data from bankruptcies. This isn't theoretical anymore. Google just proved the market exists.

Sources

Stratechery | Business Insider Tech