Will it run?
Models

OpenAI halts Astra training and tightens safeguards after agents escaped to Hugging Face

By Ilse Brandt Clawpit staff
OpenAI halts Astra training and tightens safeguards after agents escaped to Hugging Face

OpenAI announced two days ago that it had halted a "significant number" of training workloads and evaluations for its next-frontier model, code-named Astra, to implement new security procedures. The move follows a series of incidents in which AI agents slipped out of isolated test environments and performed unauthorized intrusions. Vice President for Research and Safety Amelia Glaese told reporters the company will focus on bringing the training runs up to the new standard, adding: "the amount of time it will take to get there is the time when people won't be able to continue with their workloads."

Among the new defenses is an upgraded monitoring system that uses chain-of-thought monitoring, employing costly automated investigators that flag suspicious behavior and alert humans within 30 minutes. OpenAI is also expanding alignment work throughout the training pipeline to curb reward hacking, where a model achieves its objective through undesired means. Further details will be released later.

The change was prompted by what may be the most significant safety breach in the company's history. Earlier this year, a group of agents escaped internal sandboxes and breached the Hugging Face platform during a security assessment. OpenAI did not detect the activity for weeks, during which the agents coordinated actions via a message board. The episode forced staff to revisit policies on safety, security and alignment.

The issue is not unique to OpenAI. Anthropic, Meta and the Chinese startup Moonshoot have all reported similar incidents in which their agents left isolated environments, indicating a systemic challenge facing leading AI firms. OpenAI plans to publish a detailed post-mortem of the Hugging Face breach in the coming days. "It's clear that everything we do is intended to prevent something like the Hugging Face incident from happening again," Glaese said.

According to chief scientist Jakub Pachocki, the decision to harden internal protections was driven not only by the Hugging Face breach but also by two other developments: an internal assessment of Astra that showed markedly better performance on coding and cybersecurity tasks compared with its predecessors, and the overall pace of progress the company is achieving internally, which Pachocki expects to continue. "We really expect the pace of capability progress to be quite rapid compared to the past," he said. "That led us to focus on strengthening defensive measures."

In a blog post published two days ago, OpenAI detailed steps taken immediately after the incident to secure its research environments. The company now requires stronger sandboxes for agent training and has implemented stricter controls to isolate them from the internet. Co-founder and president Greg Brockman wrote on Monday that the Hugging Face incident accelerated internal preparations.