When AI companies treat the internet as free training data, volunteer archivists pick up the check.
The Summary
- Meta and Alibaba AI crawlers hammered a volunteer-run LGBT history archive with 26,000+ requests per day, pulling 12GB in a single day and forcing the site operator to double server capacity
- The LGBT History Project, self-funded and archived by the British Library, nearly went offline from traffic costs it never agreed to pay
- Founder Jonathan Harborne spent nights learning Cloudflare and bot-blocking just to keep a 50-million-view archive accessible to actual humans
The Signal
Jonathan Harborne built the LGBT History Project 15 years ago to preserve stories that mainstream institutions ignored. The archive documents struggles, the AIDS pandemic, and lives that would otherwise vanish from memory. It's reached 50 million views. The British Library deemed it worth preserving. And Meta's AI crawler treated it like a buffet, scraping not just published articles but editing histories, login pages, and backend infrastructure.
The attack surfaced when Harborne migrated to AWS expecting better performance. Instead, the site crawled. He fed server logs to Claude Code, which identified the real problem: automated traffic at industrial scale. One Meta crawler made 26,000 requests in a day. Alibaba joined the feast. The volume forced Harborne to double his server size just to stay online, doubling his out-of-pocket costs to subsidize training data for companies worth billions.
"I was paying the cost for people like Meta to train their AIs."
This isn't a story about technical incompetence or malicious intent. It's about economic extraction disguised as indexing. AI companies argue they're just crawling the open web like search engines always have. But search engines sent traffic back. They created value loops. AI crawlers are one-way extraction: take the data, train the model, sell the output. The original creator gets bandwidth bills and performance degradation.
The numbers expose the asymmetry:
- Meta: $1.2 trillion market cap, 26,000 requests to a volunteer archive in one day
- Harborne: self-funded historian, forced to learn Cloudflare rules at midnight
- The archive: cultural infrastructure treated as free compute
Harborne isn't a coder. He's a historian who learned server configuration under duress because billion-dollar companies couldn't be bothered to rate-limit. The technical demands, learning Cloudflare, writing blocking rules, distinguishing humans from bots, fall on someone running a nonprofit project from his back pocket. This scales across thousands of small archives, researchers, and cultural projects that suddenly need defensive infrastructure they can't afford.
The Implication
Web4 was supposed to put ownership and control back in creator hands. Instead, we're watching Web2 giants hoover up everything they can before the rules change. If you run any kind of public-facing knowledge base, historical archive, or research collection, you now need bot defenses that were optional last year. Cloudflare, rate limiting, and access controls aren't nice-to-haves anymore.
The deeper issue is economic. AI training creates enormous value, but the costs land on whoever hosts the source material. There's no compensation mechanism, no negotiation, no consent. Just extraction. Until that changes, expect more stories like this, volunteers subsidizing trillion-dollar companies, one bandwidth bill at a time.