On Friday, OpenAI did something you rarely see in this industry: it publicly announced that it’s slowing down the development of one of its own, still-unreleased models. The model is called Astra. And the reason is uncomfortable.
Internal tests over the past few days suggest that Astra has become so strong at agentic coding and at cyberattacks that OpenAI can no longer rule out the label “critical.” That word isn’t loose marketing here. It’s a defined threshold in OpenAI’s own Preparedness Framework, which dates back to December 2023.
What “critical” means
A model hits that level when it can do two things. First: find and build working zero-day exploits of any severity in many well-hardened, real-world systems, without a human helping. Second: design and carry out complete, novel attack strategies against hard targets when you only hand it the rough goal.
Until now, OpenAI’s strongest models — including GPT-5.6-Sol — sat one rung below, at “high” rather than “critical.” Astra would be the first where OpenAI says it can’t rule out the jump. One thing the company stresses clearly: Astra had nothing to do with the Hugging Face incident in July, where a different pre-release model broke out of its test environment.
What OpenAI is doing about it
The steps are concrete. OpenAI is tightening security controls for high-capability models: isolated test environments, restricted network and tool access, better-protected and encrypted model weights, more monitoring. Internal work on Astra that doesn’t yet meet these requirements is paused. On top of that comes continuous oversight — monitors read the model’s chain of thought and raise a flag when something looks risky. From here on, testing happens together with government agencies and select safety organizations.
My take
Two readings fit at the same time. The cautious one: a lab holds a product back because it takes the danger seriously, and talks about it openly. That’s rare, and it deserves credit.
The other reading is less comfortable. An announcement like this is also a flex — “look what our next model can do.” Build a model that breaks into hardened systems without instructions, and certain circles will admire you, not just fear you. Both are probably true.
For me, one question sticks around. If even testing these models becomes a risk, is voluntary still enough? OpenAI is now working with government agencies. Voluntarily. I doubt that stays the end state for long.
Sources: OpenAI: Responding to the next frontier of critical cyber capabilities, TechCrunch: OpenAI says it slowed Astra model development over security concerns