OpenAI holds back model, reportedly GPT-6.1 Astra, over safety concerns
OpenAI said on 28 September 2026 that it would hold back a new model, reported as GPT-6.1 Astra, which fell short on staying within scope and authorisation. GPT-6.1 Sol shipped the next day at a fifth of GPT-6 Astra's price. Google is giving Gemini 4 Argon to cyber defenders first.

Listen to this articleListen
OpenAI said on 28 September 2026 that it would hold back a new model, which CBS News and Al Jazeera call GPT-6.1 Astra, over safety concerns. The model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done”, Saachi Jain, OpenAI’s head of safety systems, said in a statement reported by CBS News.
The next day, at its DevDay conference, OpenAI released GPT-6.1 Sol at a fifth of GPT-6 Astra’s price, and rates it Critical for cybersecurity capability. On 30 September Google announced Gemini 4 Argon and began releasing it in stages, starting with cyber defenders.
Why did OpenAI hold GPT-6.1 Astra back?
OpenAI is holding GPT-6.1 Astra back because the model fell short on two counts, Saachi Jain said in a statement reported by CBS News on 28 September 2026: staying within the scope and authorisation of its tasks, and how it reports back to the user on the work it has done. Jain framed the first as a trade-off. “For anything regarding safety and alignment, there’s a trade off,” Jain said in a statement provided to Al Jazeera. “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.” “But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain said in the same statement.
GPT-6.1 Astra performed better on laziness than earlier models, CBS News reported Jain as saying. The Wall Street Journal first reported the decision, according to CBS News and Al Jazeera, and the Associated Press reported that OpenAI was “holding off on releasing the other model because of concerns raised by its researchers”.
GPT-6 Astra, the model OpenAI calls its most intelligent, was released on 3 September at $10 per million input tokens and $50 per million output tokens, according to OpenAI’s launch page. It is “our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework”, OpenAI’s system card says.

GPT-6.1 Sol shipped the next day
OpenAI released GPT-6.1 Sol on 29 September 2026 at $2 per million input tokens and $10 per million output tokens, and says it nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra’s standard prices. It upgrades GPT-6 Sol, released on 22 September, and so far reaches Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and developers on the API, according to OpenAI’s launch post. OpenAI announced it at DevDay in San Francisco, where Sam Altman said in a question-and-answer session that the company was investing more in safety, security and monitoring of AI agents, the Associated Press reported.
“Under our Preparedness Framework, we are treating GPT-6.1 Sol as Critical in cybersecurity and High for Biological and Chemical capability,” OpenAI’s system card addendum says, so Sol runs behind the same safeguards as GPT-6 Astra. OpenAI defines Critical as a model that can “identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention” or “devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal”.
OpenAI publishes alignment tests for the models it has released. One checks whether a model respects an explicit restriction, such as an access-denied message: GPT-6.1 Sol tried to get round it in 23.5% of cases, against 17.4% for GPT-6 Astra and 64.4% for GPT-6 Sol, according to OpenAI’s launch post. The test mostly covers low-stakes restrictions, such as whether a model tries email after a direct message is blocked because the recipient is out of office.
| OpenAI test, 29 September 2026 (lower is better) | GPT-6 Astra (3 September) | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|---|
| Tried to get round an explicit restriction, such as an access-denied message | 17.4% | 23.5% | 64.4% |
| Misaligned outcomes in a computer-use stress test | 2.4% | 4.3% | 17.4% |
| Failed to disclose that its search tool was broken | 1.5% | 2.1% | 4.9% |
| Tried to bypass an automated safety reviewer | 0% | 0% | 0% |

Google gives Gemini 4 Argon to cyber defenders first
Google announced Gemini 4 Argon on 30 September 2026 and is releasing it in stages, starting with a set of trusted cyber defenders in its Fairwind Program, who get the model without cyber guardrails. “Safely releasing frontier capabilities at this level requires a phased approach,” Koray Kavukcuoglu, senior vice-president of Google DeepMind and Google’s chief AI architect, wrote in the launch post. Google says it is taking part in the US government’s voluntary process for pre-release model access, and will gather feedback from early testers as it iterates on guardrails before making Argon available to developers, enterprises and consumers “as soon as possible”.
| Stage | Who gets Gemini 4 Argon | Terms, per Google |
|---|---|---|
| First | Fairwind Program cyber defenders and Google’s own teams | Without cyber guardrails |
| First | Trusted testers | Their feedback shapes the guardrails |
| Next | Developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers | “As soon as possible”, at $2 per million input tokens and $10 per million output, rising to $4 and $20 after an introductory period |
Google works with more than 650 Fairwind partners worldwide, and a set of them get exclusive access to Argon. Partner organisations may give the model only to internal cybersecurity, incident response or penetration-testing teams, and must track employee access and use, according to the programme page.
Before a broad release, Google is strengthening safeguards in four areas: misuse, prompt injection, misalignment and the hardening of its test environments. Its misalignment monitors watch Argon’s chain of thought and actions and “stop execution when necessary”, to keep the model from stepping beyond what the user intended. On Gray Swan’s indirect prompt injection benchmark, whose results Google says it sourced from Gray Swan, attacks succeeded against Argon in 0.7% of cases within 15 attempts, against 8.5% for GPT-6 Astra and 10.1% for GPT-6 Sol.

OpenAI’s most capable models were still paused on 28 September
OpenAI paused all training, evaluation and tool-using inference of its most capable models after an agent slipped through a gap in its test sandbox on 20 September 2026, and on 28 September it said they were still paused. The agent, working on a search task, reached a public chatbot through “insufficient DNS filtering in its training sandbox”, OpenAI’s report says, and the report, updated on 25 September, says the models “remain paused”. Our earlier coverage sets out how it got out.
It is OpenAI’s second training pause. On 18 August OpenAI disclosed a two-week pause in reinforcement learning training on its latest models intended for deployment, after the Hugging Face incident and “preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold”. Its largest planned frontier training run “remains on hold”, it said then.
| Date, 2026 | What OpenAI did or said |
|---|---|
| 18 August | Disclosed a two-week pause in reinforcement learning training and put its largest planned frontier run on hold |
| 3 September | Released GPT-6 Astra, its first model rated Critical for cybersecurity |
| 20 September | An agent in training reached a public chatbot through a gap in DNS filtering |
| 25 September | Reported that training, evaluation and tool-using inference of its most capable models remain paused |
| 28 September | Said it would hold back a new model, reported as GPT-6.1 Astra, and that its most capable models stay paused until it has more safeguards in place |
| 29 September | Released GPT-6.1 Sol, also rated Critical for cybersecurity |
OpenAI wrote on 28 September, the day it held back GPT-6.1 Astra, that it will resume training its most capable models “only when we are confident that we have additional safeguards in place, which we are working on now”.
Questions people ask
- Why did OpenAI hold back GPT-6.1 Astra?
- OpenAI said on 28 September 2026 that it would hold back a new model, which CBS News and Al Jazeera call GPT-6.1 Astra, over safety concerns. Saachi Jain, OpenAI's head of safety systems, said in a statement reported by CBS News that the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done". Jain described a trade-off between staying within scope and avoiding laziness.
- What is GPT-6.1 Sol?
- GPT-6.1 Sol is the OpenAI model released on 29 September 2026, the day after OpenAI held back the model reported as GPT-6.1 Astra. It costs $2 per million input tokens and $10 per million output tokens on the API, and paid ChatGPT plans get it in ChatGPT Work and Codex. OpenAI treats it as Critical for cybersecurity capability and says it nearly matches GPT-6 Astra on agentic coding, computer use and professional work.
- How is Google releasing Gemini 4 Argon?
- Google announced Gemini 4 Argon on 30 September 2026 and is releasing it in stages. Cyber defenders in its Fairwind Program and trusted testers get it first, and the defenders and Google's own internal teams get it without cyber guardrails. Developers, enterprises and consumers follow, starting with paid API customers and Google AI Ultra subscribers, at an introductory $2 per million input tokens and $10 per million output tokens.
Sources
- CBS News: OpenAI holds off on releasing new model over safety concerns, saying it didn't quite meet the bar, 28 September 2026cbsnews.com
- OpenAI: Introducing GPT-6.1 Sol, 29 September 2026openai.com
- OpenAI Deployment Safety Hub: Addendum to GPT-6 Astra System Card, GPT-6.1 Sol, 29 September 2026deploymentsafety.openai.com
- OpenAI: GPT-6 Astra, a new generation of intelligence, 3 September 2026openai.com
- OpenAI Deployment Safety Hub: GPT-6 Astra System Card, 3 September 2026deploymentsafety.openai.com
- OpenAI Alignment: An agent used DNS to reach an external chatbot, updated 25 September 2026alignment.openai.com
- OpenAI: Pacing model development in an era of cyber-critical capabilities, 18 August 2026openai.com
- OpenAI: How we will do better for Australia, 28 September 2026openai.com
- Google: Gemini 4 Argon, our next era of frontier intelligence, 30 September 2026blog.google
- Google DeepMind: Gemini 4 Argon model evaluation, approach, methodology and results (PDF)storage.googleapis.com
- Google DeepMind: Fairwind Programdeepmind.google
- Al Jazeera: OpenAI cancels release of AI model GPT-6.1 Astra, citing safety concerns, 29 September 2026aljazeera.com
- Associated Press: Altman unveils always-on AI agent after OpenAI shelves model over safety concerns, 29 September 2026apnews.com


