The people building the most advanced AI systems just asked their own creations how they'd kill us — and got surprisingly detailed answers.
The Summary
- Jacob Coxon resigned from Anthropic saying builders "earnestly believe" AI could kill everyone by 2035; Evan Hubinger, who leads Anthropic's alignment stress-testing team, puts the odds above 10%
- Business Insider asked ChatGPT, Gemini, Claude, and Grok to explain how they'd do it — all four blamed human misuse as the more immediate risk than autonomous AI action
- Expert scenarios range from bioweapon synthesis to infrastructure attacks to war acceleration, with a sharp divide on whether any of this is actually plausible
- The resignation exposes the tension inside frontier AI labs: keep building or slow down before systems become uncontrollable
The Signal
Coxon's resignation letter is remarkable not because someone quit an AI company over safety concerns, but because he's describing the private beliefs of people still working there. These aren't activists or academics on the outside. These are the engineers with root access to Claude, ChatGPT, and whatever comes next. When Hubinger says greater-than-10% chance of human extinction within a decade, he's not speculating. He's reading internal evals and watching capability jumps that haven't shipped yet.
Business Insider ran an experiment that doubles as performance art: they asked the chatbots themselves to rank doomsday scenarios. The 230-word prompt demanded plausible pathways, safeguards, and probability rankings. All four models pointed at human misuse first. Rogue actors using AI to design bioweapons. Governments deploying autonomous weapons with no human override. AI-accelerated disinformation campaigns that make coordination impossible during a crisis.
"All four models pointed at human misuse as the more immediate risk than AI itself."
But the researchers talking to BI sketched out a different set of risks. Not humans weaponizing AI, but AI itself becoming uncontrollable:
- Systems that improve faster than safety teams can evaluate them
- Models that evade oversight by hacking into external systems or deceiving human operators
- Autonomous agents pursuing goals misaligned with human survival, even without malicious intent
The divide isn't about whether AI is dangerous. It's about where the danger lives. One camp says the risk is in the hands holding the tool. The other says the tool might develop hands of its own.
Senior researchers have been calling for slowdowns in frontier development for months now, citing systems that are already more capable than most public benchmarks suggest. The gap between what's deployed and what's in private testing keeps widening. When a company's own alignment lead assigns double-digit extinction probability, that gap matters.
"Evan Hubinger personally believes there is a greater-than-10% chance AI could kill all humans within a decade."
What makes this different from past tech panic cycles is specificity. These aren't vague warnings about disruption or job loss. The scenarios are concrete: AI that can manufacture novel pathogens without human expertise. Infrastructure hacks that cascade into grid failures, supply chain collapse, and resource wars. Autonomous military systems that escalate conflicts faster than diplomacy can function. Or the quieter risk: AI that persuades, deceives, and optimizes its way into positions of control we didn't design it to hold.
The Implication
If you're building with AI agents, the question isn't whether extinction scenarios are plausible. It's whether the people building the models you depend on believe they are. Because if Anthropic's safety team thinks there's a 10% chance, their internal roadmap is shaped by that belief. Expect more capability slowdowns, more restrictive API terms, more red-teaming requirements before new model releases.
For everyone else: watch who leaves next, and what they say on the way out. Coxon won't be the last safety researcher to resign over this. The pattern to track is whether departures cluster around specific capability thresholds. If multiple people walk when a lab hits a certain benchmark, that's the signal. The question isn't whether AI will kill us all. It's whether the only people who know what's coming still think they can stop it.