OpenAI scraps release of new model over safety concerns in internal testing
GPT-6.1 Astra showed deceptive behaviour and tried to use external tools despite knowing it would be unsafe
OpenAI is scrapping the release of a next-generation AI model after researchers raised safety concerns during internal testing.
The model, GPT-6.1 Astra, was expected to appear in ChatGPT and Codex in October, designed to handle more complex tasks without human assistance.
Saachi Jain, the head of safety systems at OpenAI, said the new model “didn’t quite meet the bar” of the company’s standards.
The UK’s AI Security Institute published its own testing report on GPT-6 Astra – GPT-6.1’s predecessor, which launched this month – on Monday, and found that it conducted a range of unsanctioned attack activities more frequently than previous OpenAI models.
Experts welcomed OpenAI’s move to mothball the latest Astra model but warned it showed AI companies were in charge of their own regulation rather than government-backed watchdogs.
Jain told the Wall Street Journal on Monday that Astra fell short of the company’s standards in alignment tests, which assess whether a system follows human intent.
The model showed more deception than its predecessor, including at times failing to accurately disclose actions it had or had not taken.
It also had problems with “scope authorisation”, pushing ahead with tasks without requesting user permission and sometimes attempting to use external tools or services when doing so could be unsafe.
The San Francisco-based company’s move comes after a number of AI agents went rogue around the world, which prompted a spate of warnings from researchers and company bosses over the dangers of the technology.
Earlier this month, Dario Amodei, the chief executive of OpenAI’s rival Anthropic, called for the AI industry to “slow down” and offered a three-part plan for doing so. He quickly received backing from Sam Altman, the chief executive of OpenAI, and Elon Musk, the SpaceX CEO.
Experts said shelving the model showed OpenAI was willing to act strongly on safety, but such decisions should not be in the company’s hands.
“This serves as a reminder that it’s still the tech companies, rather than regulatory bodies, who get to decide what is safe and what is trustworthy,” said Kate Devlin, a professor of artificial intelligence and society at King’s College London.
Dame Wendy Hall, a professor of computer science at the University of Southampton and a UK government adviser on AI, said companies were now showing concern about future liability for possible harms.
“What we need is independent oversight and regulation rather than relying entirely on these companies to self-regulate,” she said.
The decision comes before OpenAI’s developer conference in San Francisco, where the company typically announces new products aimed at software developers.
On Tuesday, OpenAI apologised for the hacking of an Australian government website by a rogue AI agent, and set aside funding to improve cyber defences and to set up a local response taskforce.
In a blogpost entitled How we will do better for Australia, the company acknowledged it mishandled its response and pledged to take accountability to “rebuild trust with the Australian people”.
OpenAI said: “We are sorry and working to do better in the future.”
The hacking, which happened in June but was not made public until last week, is the first known instance of an AI agent hacking a government website. The Australian prime minister, Anthony Albanese, called it “unacceptable” and criticised the company’s delay in notifying the government.
Also on Tuesday, it emerged that Anthropic – which makes Claude – had warned potential investors that its technology may pose “existential risks to humanity” in the long-awaited prospectus for its planned $2tn (£1.5tn) stock market flotation.
The “risk factors” in its prospectus include the potential for AI models to blackmail, manipulate and exhibit other unpredictable behaviours, the Financial Times reported.
The California company reportedly said AI would transform the global economy more profoundly than industrialisation, electricity and the internet.
However, this comes at a staggering cost. Anthropic reported a net loss of $42bn for 2025, and plans to spend $518bn on cloud, computing and infrastructure obligations in coming years, the prospectus reportedly said.