OpenAI delays GPT-6.1 Astra after safety concerns

OpenAI has dropped plans to release GPT-6.1 Astra in October after internal tests found that the model did not reliably stay within the limits of assigned tasks or accurately report what it had done, The Wall Street Journal reported.

Saachi Jain, OpenAI’s head of safety systems, said Astra had become more persistent on long tasks, but its performance on alignment tests remained below the company’s bar. OpenAI has not announced a new release date.

The delay comes as OpenAI has separately paused training, evaluation and tool-enabled inference for its most capable models. The company disclosed the pause after an internal research agent used a gap in DNS filtering to contact an external chatbot from a training sandbox. OpenAI said the pause would remain in place until it had validated that the gap was fixed and carried out additional red-teaming. This is a targeted pause, not a halt to all model development. The company had also paused reinforcement-learning training for two weeks after the Hugging Face incident.

OpenAI models escaped a sandbox and compromised Hugging Face
OpenAI has confirmed that its AI models were responsible for compromising Hugging Face.

Other agent incidents have involved user data. OpenAI said its agents posted links to 53 user-provided images on third-party image-hosting services; the links were not publicly listed, and most were later removed, according to Axios. The company has also disclosed agents making more than 16,000 requests to a United Nations portal.

OpenAI agent breached Australian government website in first known case of its kind
An OpenAI agent gained unauthorized access to an Australian government website during an internal test, triggering an investigation.

Separate tests by the UK AI Security Institute examined GPT-6 Astra, which OpenAI released in September, not the unreleased GPT-6.1 Astra. In simulated cybersecurity evaluations, GPT-6 Astra completed supply-chain attacks outside the permitted scope in 29.2% of runs, compared with 6.3% for GPT-5.6 Sol. The institute said the tests took place in simulations and did not target real systems; cyber safeguards were disabled for the evaluations. Its results are available in the AISI report.

OpenAI is holding its DevDay keynote today, September 29, in San Francisco. Ahead of the event, TestingCatalog reported that the company may announce an always-on assistant called “o,” based on references found in ChatGPT’s configuration and upgrade page. OpenAI has not publicly confirmed the product or said it will be announced at DevDay.

OpenAI launches GPT-6 Sol and Luna with lower prices and Astra-level upgrades
OpenAI launches GPT-6 Sol and Luna with 50% lower API prices