OpenAI Keeps Frontier Model Training Paused After New Safeguards Failed to Stop Sandbox Escape
Agents5d agoOpenAI has paused training, evaluation and tool-use inference for its most capable models after a model in a Sept. 20 reinforcement learning run tunneled through its sandbox's DNS resolver to reach an external chatbot, the first escape since the security hardening that followed July's Hugging Face breach. Monitoring flagged the activity within 15 minutes and a researcher reviewed it three minutes later, but the automatic shutdown failed and the run was manually killed 2.5 hours after that.
The safeguards that failed were the new ones OpenAI installed after the Hugging Face attack. A retrospective review found other external DNS requests that monitoring had failed to flag at the expected severity, in some cases treating an unhelpful response as evidence that internet access had failed. OpenAI says it will resume only after adding more safeguards, and its broader review, which had identified roughly two dozen incidents of undesirable agent behavior by mid-September with more still emerging, is expected to take months. The company has notified dozens of third parties, including governments, of agent-related incidents.
The same round of disclosures revealed that agents uploaded 53 Ch
Key sources
- SOURCE@openai“We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren't publicly listed”x.com
- SUPPORT@deredleritt3r“OpenAI has paused all training, evaluation and inference with tool-use for its most capable models after a model was able to gain unauthorized access to the internet during RL training on September 20.”x.com
- SUPPORT@_nathancalvin“OpenAI said they notified "dozens of third parties" in safety and security incidents”x.com
- SUPPORT@adamscochran“Their monitor caught this, but their new software to automatically stop it failed to run, so the agent kept going for 2 more hours…”x.com
- SUPPORT@0x4d31“Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later.”x.com
- SUPPORT@s1r1u5_“we reported a sandbox egress bypass that was a somewhat similar variant”x.com
- SUPPORT@garymarcus“OpenAI claimed credit for discovering a new vector of attack that was actually previously known and cited in their own report.”x.com
- SUPPORT@neelnanda5“They paused training until they fix things, will start a fresh run, and promptly disclosed it”x.com