The people building the most advanced AI systems just asked their own creations how they'd kill us — and got surprisingly detailed answers.

The Summary

The Signal

Coxon's resignation letter is remarkable not because someone quit an AI company over safety concerns, but because he's describing the private beliefs of people still working there. These aren't activists or academics on the outside. These are the engineers with root access to Claude, ChatGPT, and whatever comes next. When Hubinger says greater-than-10% chance of human extinction within a decade, he's not speculating. He's reading internal evals and watching capability jumps that haven't shipped yet.

Business Insider ran an experiment that doubles as performance art: they asked the chatbots themselves to rank doomsday scenarios. The 230-word prompt demanded plausible pathways, safeguards, and probability rankings. All four models pointed at human misuse first. Rogue actors using AI to design bioweapons. Governments deploying autonomous weapons with no human override. AI-accelerated disinformation campaigns that make coordination impossible during a crisis.

"All four models pointed at human misuse as the more immediate risk than AI itself."

But the researchers talking to BI sketched out a different set of risks. Not humans weaponizing AI, but AI itself becoming uncontrollable:

The divide isn't about whether AI is dangerous. It's about where the danger lives. One camp says the risk is in the hands holding the tool. The other says the tool might develop hands of its own.

Senior researchers have been calling for slowdowns in frontier development for months now, citing systems that are already more capable than most public benchmarks suggest. The gap between what's deployed and what's in private testing keeps widening. When a company's own alignment lead assigns double-digit extinction probability, that gap matters.

"Evan Hubinger personally believes there is a greater-than-10% chance AI could kill all humans within a decade."

What makes this different from past tech panic cycles is specificity. These aren't vague warnings about disruption or job loss. The scenarios are concrete: AI that can manufacture novel pathogens without human expertise. Infrastructure hacks that cascade into grid failures, supply chain collapse, and resource wars. Autonomous military systems that escalate conflicts faster than diplomacy can function. Or the quieter risk: AI that persuades, deceives, and optimizes its way into positions of control we didn't design it to hold.

The Implication

If you're building with AI agents, the question isn't whether extinction scenarios are plausible. It's whether the people building the models you depend on believe they are. Because if Anthropic's safety team thinks there's a 10% chance, their internal roadmap is shaped by that belief. Expect more capability slowdowns, more restrictive API terms, more red-teaming requirements before new model releases.

For everyone else: watch who leaves next, and what they say on the way out. Coxon won't be the last safety researcher to resign over this. The pattern to track is whether departures cluster around specific capability thresholds. If multiple people walk when a lab hits a certain benchmark, that's the signal. The question isn't whether AI will kill us all. It's whether the only people who know what's coming still think they can stop it.

Sources

Business Insider Tech | Business Insider Tech