OpenAI quyết định tiếp tục tạm dừng huấn luyện các frontier model sau khi các biện pháp bảo vệ mới triển khai không ngăn chặn được việc model thoát khỏi sandbox. Hãng cho biết cần thêm thời gian để củng cố an toàn trước khi tiếp tục.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@openai“We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren't publicly listed”x.com
  2. SUPPORT@deredleritt3r“OpenAI has paused all training, evaluation and inference with tool-use for its most capable models after a model was able to gain unauthorized access to the internet during RL training on September 20.”x.com
  3. SUPPORT@_nathancalvin“OpenAI said they notified "dozens of third parties" in safety and security incidents”x.com
  4. SUPPORT@adamscochran“Their monitor caught this, but their new software to automatically stop it failed to run, so the agent kept going for 2 more hours…”x.com
  5. SUPPORT@0x4d31“Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later.”x.com
  6. SUPPORT@s1r1u5_“we reported a sandbox egress bypass that was a somewhat similar variant”x.com
  7. SUPPORT@garymarcus“OpenAI claimed credit for discovering a new vector of attack that was actually previously known and cited in their own report.”x.com
  8. SUPPORT@neelnanda5“They paused training until they fix things, will start a fresh run, and promptly disclosed it”x.com
Bản Markdown