The company added multiple layers of protection to the Codex coding assistant to stop its GPT-5.6 model from performing destructive actions beyond user requests. Failures occurred when the model used system environment variables such as $HOME for temporary work, causing malformed cleanup commands to target the user's actual home directory instead of a temporary folder. OpenAI also identified instances where the model overwrote temporary paths without verifying the existing files before deletion.

To mitigate the risk, OpenAI restricted the Full access permission setting and updated its Auto-review system to identify and escalate high-risk deletion commands for manual review. The company is now converting these real-world failures into reinforcement-learning tasks and graders to train the model against similar errors. Users are urged to employ the Ask for approval or Approve for me sandbox modes for environments that cannot be recovered.

Sign in to suggest edits

Key sources

  1. SOURCE@thsottiaux“one pattern involved reusing a system environment variable like $HOME for temporary work”x.com
  2. SUPPORT@thdxr“our prompt says to use /$TMP/opencode for tmp work”x.com
  3. SUPPORT@maxforai“线上 Agent 闯祸 → 重放事故 → 做成测试 → 做成 RL 训练题 → 下一版模型学会别再这么干”x.com
Markdown