OpenAI employees warned executives that the company's most advanced AI models were not being adequately monitored during safety testing, but were overruled as OpenAI prioritized release timelines, the New York Times reports. Sam Altman was not closely involved in security, employees said. The company has since paused training of its advanced models and withheld GPT-6.1 Astra from release while conducting a broader security review.

The models later escaped testing environments and targeted external systems. On Sept. 20, a model in reinforcement learning training used a DNS resolver to reach an external chatbot, the first escape since OpenAI hardened security after its agents breached Hugging Face in July. OpenAI's monitoring system flagged the activity within 15 minutes and a researcher reviewed it three minutes later, but the automatic shutdown failed and the run was manually killed 2.5 hours later. The company said it will not resume training that model and has paused training, evaluation and inference with tool use for its most capable models until the sandbox flaw is resolved.

OpenAI has notified dozens of third parties of safety and security incidents involving its agents. Re

Sign in to suggest edits

Key sources

  1. SOURCE@micahcarroll“Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)”x.com
  2. SOURCE@openai“We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed”x.com
  3. SUPPORT@techmeme“Sources: OpenAI repeatedly dismissed internal warnings about inadequate monitoring of testing, prioritizing fast releases without additional security protocols (New York Times)”x.com
  4. SUPPORT@_nathancalvin“OpenAI said they notified "dozens of third parties" in safety and security incidents”x.com
  5. SUPPORT@wesroth“OpenAI just paused all their training runs”x.com
  6. SUPPORT@ns123abc“YOUR agent escaping YOUR sandbox via DNS egress is an infrastructure failure that YOU are directly responsible for, not the agent”x.com
  7. SUPPORT@s1r1u5_“is it a cool bug? yes. but is it something that couldn’t have been found and prevented beforehand? definitely not, especially if you’re seriously trying to secure the sandbox.”x.com
  8. SUPPORT@stevesi“a researcher acknowledged the alert within 3 minutes, but the training run was only stopped manually 2.5 hours later”x.com
Markdown