The same week both frontier labs crossed their own red lines, they're asking you to trust the new ones.
The Summary
- OpenAI's Astra model is the first to hit "Critical" cybersecurity capability under their Preparedness Framework, triggering stronger release safeguards
- Anthropic is co-developing enterprise safeguards directly with customers, moving safety decisions out of the lab and into corporate boardrooms
- The timing is not coincidental: both companies are rewriting the playbook for what happens when models get good enough to break things
The Signal
OpenAI's Astra represents a threshold moment. Under their Preparedness Framework, models are evaluated across four risk categories: cybersecurity, CBRN (chemical, biological, radiological, nuclear), persuasion, and model autonomy. Each category has four levels: Low, Medium, High, and Critical. Astra is the first OpenAI model to reach Critical in any category. That category is cybersecurity, meaning Astra can find and exploit vulnerabilities at a level OpenAI previously said would require heightened safeguards before release.
The company is keeping those safeguards vague. "Stronger safeguards for release" is doing a lot of work in that sentence. What does stronger mean when you already had a framework? The Preparedness Framework was supposed to be the guardrail. Now we're getting guardrails for the guardrails.
"Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold, with stronger safeguards for release."
Meanwhile, Anthropic is taking a different approach with enterprise customers. Instead of setting internal thresholds and announcing when they're crossed, Anthropic is co-developing safeguards with the companies actually deploying these models. The framing is collaborative, almost consultative. We're building this together, the message says. Your risk tolerance matters.
This is either more honest or more dangerous, depending on how you squint at it. On one hand, enterprises have skin in the game. They know their threat models better than a lab in San Francisco. On the other hand, enterprise risk tolerance has given us everything from the 2008 financial crisis to the Boeing 737 MAX. Letting customers help write the safety manual for frontier AI is a bet that corporate incentives and public safety will align. History suggests they won't.
Key differences in approach:
- OpenAI: Internal framework, public thresholds, company-controlled release decisions
- Anthropic: Customer co-development, enterprise-specific safeguards, distributed decision-making
- Both: Crossing capability lines they previously drew, redefining what "safe release" means in real time
The Implication
We're watching two companies solve the same problem with opposite strategies. OpenAI is centralizing safety decisions and making them more opaque. Anthropic is distributing them and calling it collaboration. Neither approach has been tested at scale because we've never had models this capable before.
If you're building with these models, the safeguards matter less than the incentives behind them. OpenAI's framework protects OpenAI's reputation. Anthropic's customer co-development protects customer relationships. Your job is to figure out what gets protected when those priorities conflict with yours. The labs just told you their models can break critical infrastructure. Believe them. Then ask who decides when that's an acceptable risk.