YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0043

Incident record

OpenAI pauses internal work on Astra after it cannot rule out a Critical cyber capability threshold

Severity
Critical
Status
Research
Type
AI-Found Vuln
Target
OpenAI's own unreleased Astra model, under its Preparedness Framework
Actor
researcher

What happened

OpenAI said internal evaluations of Astra, an unreleased successor model, showed advances in agentic coding and cybersecurity strong enough that it cannot rule out the model reaching the Critical cyber capability threshold under its own Preparedness Framework: the level at which a model could independently find and exploit zero days in hardened real systems, or plan and run a cyberattack from a high-level goal alone. Previous models including GPT-5.6 Sol had been assessed only at the lower High threshold.

OpenAI paused internal activities involving Astra that do not yet meet strengthened security controls, such as isolated testing environments, restricted network and tool access and additional monitoring of the model's chain of thought, and said Astra was not involved in the Hugging Face breach disclosed in July. Three days later it expanded its Daybreak programme for vetted defenders and introduced GPT-5.6-Cyber, a model trained for vulnerability research with far fewer refusals than GPT-5.6 Sol. This is a self-reported capability assessment rather than a confirmed breach. On 18 August OpenAI went further: a two-week pause on reinforcement-learning training for models intended for deployment while it hardens and red-teams its research environments, its largest planned frontier RL run held, and monitoring that inspects internal model state at every sampled token, now required for all inference of Astra with tools and for all tool-using RL training and evaluation at GPT-5.6 Sol capability or above. OpenAI puts the cost of that monitoring at roughly 20 per cent of the compute being watched, on its own estimate, and gives its teams 30 minutes to alert on flagged activity and 30 minutes to rule out a false positive before the work is paused.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026