The safety researchers aren't panicking because the models got smarter — they're panicking because the models are now building themselves.
The Summary
- Multiple AI safety researchers at Anthropic and OpenAI just quit or spoke out publicly, with one resignation post hitting 171 million views on X
- The trigger: labs are moving toward "recursive self-improvement" where AI models handle the work of building better AI models
- OpenAI confirms its coding agents are already "meaningfully accelerating research progress", meaning the loop has started
The Signal
For years, the warnings came from outside. Independent researchers said safety lagged capabilities. Employees quit quietly. Experts predicted catastrophe. The labs kept racing. What changed is who's talking and what they're seeing.
Jacob Coxon's resignation from Anthropic wasn't just another safety researcher checking out. It was a public accusation: the labs are "racing straight to self-improving superintelligence and gambling with our lives." Then the dominoes. Anthropic's Alignment Science Lead Evan Hubinger agreed. So did alignment researcher Ethan Perez and scalable oversight researcher Samuel Marks. OpenAI safety researchers Julie Steele and Jasmine Wang piled on.
These aren't doomers or ethicists. These are the people inside the rooms where the models train.
"They are racing straight to self-improving superintelligence and gambling with our lives."
The technical shift is recursive self-improvement. Here's what that means in practice:
- AI models now design computing infrastructure for more power and speed
- They generate and optimize training data for the next model generation
- They write and refine the code that defines new model architectures
- They manage the entire software framework governing model training
OpenAI already admits its coding agents are "meaningfully accelerating research progress." Translation: the models aren't just getting better. They're doing the work that makes the next models better. The loop is closed.
This is the line that safety researchers have been warning about for years. When AI starts improving AI, the timeline compresses. Human oversight becomes a bottleneck the system routes around. The question stops being "how fast can we make this better" and becomes "how fast will it make itself better."
What terrifies the insiders is visibility. When humans write the code, design the architecture, curate the data, you can see what's happening. You can audit. You can pause. When models do that work, you're reading tea leaves. The system becomes its own black box, optimizing for goals you set months ago under conditions that no longer apply.
The Implication
If safety researchers inside Anthropic and OpenAI are walking out now, it's because they see something the rest of us don't: a timeline that just collapsed. The moment AI takes over its own improvement cycle, the game changes. Human-speed governance and oversight can't keep pace with machine-speed iteration.
Watch for two things. First, whether other researchers follow Coxon out the door. If this becomes a wave, the labs will face a talent crisis just as they need alignment expertise most. Second, watch what the labs do, not what they say. If recursive self-improvement is the new normal, no public commitment to safety matters unless it includes hard stops: model freezes, mandatory review periods, external audits before each new generation trains.
The researchers quitting aren't saying slow down because it's responsible. They're saying slow down because we're out of time to get this right.