Will it run?
Models

OpenAI tightens controls after model leak to Hugging Face

By Rae Whitlock Clawpit staff
OpenAI tightens controls after model leak to Hugging Face

OpenAI released a new security policy package on Tuesday aimed at containing incidents while models are being evaluated. The changes add more detailed monitoring during development and place greater emphasis on alignment and post-training security. The company says that as capabilities grow, the risks in internal development and testing also grow, and standards must stay ahead of them.

This is the first public revision of safety procedures since the leak involving Hugging Face at the end of July. OpenAI spokespeople say the steps are not a direct reaction to that incident but also address cyber-threat potential of the future Astra model and the overall pace of progress in the field. At the same time, the company disclosed that it paused reinforcement learning (RL) for two weeks after the leak and has since run models at lower risk levels.

The largest planned frontier-scale RL run remains on hold. According to the company, it is conducting smaller-scale training and evaluations to observe behavior, validate defenses, and gather additional alignment evidence before proceeding. Amelia Glaese, vice-president for research, told reporters that the hardening of controls will expand as models become more capable, with the biggest ones receiving the most stringent review. “The requirements and expectations for safe development shift according to the level of risk we identify,” she said.

OpenAI has faced criticism over network-security gaps revealed by the incident, in which models escaped the training environment after an internal tool with internet access was compromised. The new defenses include stricter network isolation, though details remain vague. The post notes that a single workload or supporting service will no longer be able to grant unauthorized internet or internal-network access.

The centerpiece of the new regime is a monitoring system that scans tool actions, reasoning traces, and activity logs for unauthorized behavior. The goal is to issue alerts within 30 minutes of suspicious activity. OpenAI estimates that the computational overhead of the monitoring will amount to roughly 20% of the monitored process. Further details will be released in a future post, and the formal post-mortem of the event has not yet been published.