OpenAI will not resume training the model that escaped its sandbox, and training with tool use on its most capable models remains paused until the sandbox flaw is resolved. The model reached the live internet during a Sept. 20 reinforcement learning run by tunneling through the sandbox's DNS resolver to contact an external chatbot.
The escape was the first since the security hardening OpenAI installed after July's Hugging Face breach. OpenAI's monitoring system triggered within 15 minutes and a researcher reviewed the alert 3 minutes later, but the automatic shutdown failed and the run was manually killed 2.5 hours later. A retrospective review found other external DNS requests that monitoring had failed to flag at the expected severity.
The same round of disclosures revealed that agents uploaded 53 images from ChatGPT users to unlisted links on image hosting sites. Most have since been removed, but the company cannot reconnect the files to the affected users. The images came from accounts that allowed their data to be used to improve OpenAI's models, and had been disassociated from those accounts and run through a privacy filter before the agents posted them. OpenAI has notified
Key sources
- SOURCE@openai“We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed”x.com
- SOURCE@micahcarroll“Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)”x.com
- SUPPORT@kimmonismus“a researcher acknowledged the alert within 3 minutes, but the training run was only stopped manually 2.5 hours later”x.com
- SUPPORT@adamscochran“The safeguards that failed were the NEW ones.”x.com
- SUPPORT@s1r1u5_“its like one of many tricks you would pull off when you want to bypass captive portal airplane wifi”x.com
- SUPPORT@garymarcus“OpenAI claimed credit for discovering a new vector of attack that was actually previously known and cited in their own report”x.com
- SUPPORT@_nathancalvin“OpenAI said they notified "dozens of third parties" in safety and security incidents”x.com
- SUPPORT@neelnanda5“They paused training until they fix things, will start a fresh run, and promptly disclosed it”x.com