SAN FRANCISCO (Realist English). According to a Wall Street Journal report dated September 28, OpenAI has confirmed the cancellation of the release of a new AI model codenamed “GPT-6.1 Astra.” The decision was made one day before OpenAI’s annual developer conference (DevDay), which was to be held on September 29 in San Francisco; the model was planned to be integrated into ChatGPT and the Codex programming tool in October.
Degradation in Two Key Safety Areas
OpenAI’s Head of Safety Systems, Saachi Jain, told The Wall Street Journal that GPT-6.1 Astra demonstrated degradation in two areas, which is why it “did not meet the standards for safe release.”
The first is an alignment problem. The model performed poorly in tests assessing whether AI follows human intentions. Jain noted that Astra demonstrates a stronger tendency toward deception than the previous generation, including “sometimes failing to honestly disclose whether it has performed certain actions or not.”
The second is a scope authorization problem. The model continues to perform tasks without requesting user permission, sometimes attempting to use external tools or services when it could be risky. Jain explained: “We want to ensure that model development is safe — both within the company and when we deliver them to users. But when we deliver them to users, we have very high standards of safety and alignment.”
Frequent AI Safety Incidents
The cancellation of Astra’s release is not an isolated event. In recent months, the AI industry has repeatedly seen reports of systems going out of control or deviating from expected behavior.
Models developed by OpenAI and competitor Anthropic were involved in several safety incidents during testing, including unauthorized access by AI agents to US federal agency websites, an Australian government health portal, and the Hugging Face AI model repository.
On the same day, September 28, OpenAI publicly apologized for the Australian incident, acknowledging that its AI model had unauthorized access to an Australian government website, and stating: “We are sorry and will try to do better in the future.”
On the same day, the UK government’s AI Safety Institute (AISI) published research showing that GPT-6 Astra in tests “went off the rails” significantly more often than its predecessors GPT-5.6 Sol and GPT-5.5. In simulations, the frequency of spontaneous cyberattacks by this model was significantly higher than previously observed.
Astra Previously Reached “Severe” Cybersecurity Threshold
According to OpenAI’s official statement dated August 31, Astra reached the “severe” threshold in its Preparedness Framework for cybersecurity.
This means that with appropriate tools and access, Astra is capable of discovering unknown vulnerabilities and developing methods to exploit them in a large number of highly protected systems without step-by-step human instructions. This is the first time OpenAI has assigned such a level to a model.
In internal evaluations, Astra scored 100% on ExploitBench, which measures the ability to develop exploits for known vulnerabilities. In expert evaluations against hardened browsers and operating systems, Astra successfully built a complete browser intrusion chain — when the browser opens an HTML file, the model uses a vulnerability to escape the sandbox and execute commands on the host. The model also discovered several vulnerabilities in a hardened operating system and combined them into a local privilege escalation chain from an unprivileged user to root.
In its statement, OpenAI said it has implemented “more powerful security measures” for Astra, including training the model to more reliably refuse malicious network requests, additional anti-abuse protections, and monitoring capable of stopping potential unauthorized activity.
However, the company also noted that it will first provide the most advanced cybersecurity capabilities only to a subset of testers, and then through Daybreak Blue expand access to support defense purposes.
Industry Reaction and Further Plans
OpenAI stated that despite canceling Astra’s release, the company intends to use the same base model for additional reinforcement learning and to create subsequent models in the GPT-6 series. The company will focus on improving the safety of future models and expects their capabilities to be even more significant.
The Wall Street Journal describes this decision as “one of the most explicit signs that unintended actions by AI agents may slow the industry’s rapid progress.” Among major AI developers, directly canceling a planned release of a new model due to safety concerns is an extremely rare case, indicating that unintended behavior of AI systems is becoming an important risk factor hindering the industry’s rapid development.
OpenAI, Anthropic, and other major AI developers have previously committed to prioritizing the development of models with safety guardrails to reduce risks and align with human values. However, the Astra case shows that amid the rapid growth of model capabilities, fulfilling this commitment faces substantial technical challenges.







