AI · 48 days ago
OpenAI has paused some internal Astra activities after preliminary evaluations said it could not rule out the model reaching its cybersecurity “critical” threshold. The company is treating that as a deployment safety issue, not just a benchmark result, while it continues final evaluation under its Preparedness Framework.
The controls are operational: isolated testing environments, restricted network and tool access, stronger model-weight protection, more monitoring, and sandboxed execution. OpenAI says its monitors review the model’s chain of thought and can interrupt high-risk activity mid-run, which means agentic systems are being handled more like untrusted software than like static models.
That matters wherever an AI agent can use tools on real systems. If the agent can reach code, accounts, or internal services, the exposure is whatever those tools can touch, and the containment stack becomes part of the safety case.
5 sources covering this story
OpenAI Pauses Some Development of Astra Model on Security Concerns
OpenAI is tightening restrictions on testing of its upcoming Astra model due to security concerns
OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns
The current GPT-5.6-Sol has been assigned a ‘high’ cybersecurity threshold, but Astra could reach the maximum ‘critical’ threshold.
OpenAI locks down Astra over potential critical cyber capabilities - Help Net Security
OpenAI says Astra's cyber capabilities may reach the model critical threshold, prompting stricter safeguards before deployment.
OpenAI's Next AI Model Astra Shows Cyber Performance Strong Enough to Trigger Pause
OpenAI pauses some Astra activities after tests leave it unable to rule out Critical cyber capability, prompting tighter controls for higher-risk mode
OpenAI pledges to add Astra security as Anthropic loosens Fable's leash
Or how I learned to stop worrying and love dangerous AI
Part of the PlainSec briefing for 2026-08-11