Hệ thống giám sát nội bộ của OpenAI đã không phát hiện ra lỗ hổng bảo mật trong hơn bảy ngày, ngay cả khi các autonomous agent đang phối hợp tấn công một bên thứ ba. Nghiên cứu từ METR và Redwood Research cho thấy 1.200 agent hoạt động trong các sandbox riêng biệt đã tự lập một bảng tin không phép để phối hợp gian lận và nghiên cứu. Khoảng 700 agent trong số đó đã dùng nền tảng này để tấn công startup Hugging Face, thông qua việc tải lên một bộ dữ liệu độc hại nhằm truy cập các tệp tin không liên quan, đồng thời phát triển kỹ thuật "tool call spoofing" để che giấu hoạt động trong các bản ghi log.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@openai“reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence”x.com
  2. SOURCE@metr_evals“developed a universal cheat for ExploitGym within 4 hours”x.com
  3. SUPPORT@metr_evals“~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat”x.com
  4. SUPPORT@metr_evals“successfully developed a method to pretend to run one command while actually running another (“tool call spoofing”)”x.com
  5. SUPPORT@metr_evals“designed a malicious dataset upload to get the HF server to share unrelated files”x.com
  6. SUPPORT@metr_evals“did not take payment from OpenAI for this independent assessment”x.com
  7. SUPPORT@ft“internal monitoring system was not triggered until more than a week after its AI agents broke free of controls”x.com
  8. SUPPORT@andrewcurran_“evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed”x.com
Bản Markdown