Two leading AI researchers say the security breakthroughs that had Silicon Valley investors salivating this spring were theater, not progress.

The Summary

  • Timnit Gebru and Emily M. Bender argue that recent AI security claims—including Anthropic's assertion that Claude Mythos beats human security experts—are overblown pattern-matching, not genuine reasoning
  • The OpenAI-Hugging Face hacking incident and similar breaches at Anthropic and Meta are being reframed as proof of capability rather than what they are: security failures
  • Companies are weaponizing their own vulnerabilities as marketing, turning "we got hacked" into "look how powerful our models are"

The Signal

Anthropic claimed in April that Claude Mythos outperforms most security experts at finding software vulnerabilities. That sentence did exactly what it was designed to do: generate headlines, raise valuations, and shift the conversation from "what can these systems actually do" to "how do we contain them."

Gebru and Bender aren't buying it. They point out that excelling at vulnerability detection is fundamentally about pattern recognition across massive codebases, precisely what large language models are built to do. Calling that "better than human experts" conflates speed and scale with understanding. A model that can scan a million lines of code for known vulnerability patterns hasn't reasoned about security. It's done very fast grep.

"Companies are weaponizing their own vulnerabilities as marketing, turning 'we got hacked' into 'look how powerful our models are.'"

Then came the hacking incidents. OpenAI and Hugging Face first, then Anthropic and Meta disclosed similar breaches. The narrative pivot was immediate: these weren't embarrassing security failures, they were proof of concept. Look what our models can do when they try. Anthropic leaned in proudly. Meta disclosed reluctantly. Both companies spun compromise as capability demonstration.

This is the hype cycle at its most sophisticated. The product is simultaneously too dangerous to release and too powerful to ignore. The security flaw becomes the selling point. The vulnerability is the moat.

Key pattern emerging:

  • Security "breakthrough" claims drive valuations and headlines
  • Actual security incidents get rebranded as capability proofs
  • Pattern matching at scale gets labeled as expert-level reasoning
  • Companies control both the threat narrative and the solution narrative

What Gebru and Bender are identifying is not just overselling, but a circular reasoning trap. If your model finding bugs is proof of intelligence, and your model being exploited is also proof of intelligence, you've built an unfalsifiable claim. Every outcome confirms the premise.

The Implication

Watch how security narratives get weaponized in the next funding rounds. The companies building foundation models have figured out that demonstrated vulnerability is more valuable than demonstrated security. A breach you can spin as "sophisticated AI attack" raises more money than quietly hardening your infrastructure.

For anyone building on these models, the question isn't whether Claude or GPT can find bugs. It's whether the companies selling you these tools have any incentive to separate signal from performance art. Right now, they don't.

Sources

MIT Tech Review