Trong báo cáo kỹ thuật, OpenAI nhận định sự cố tại Hugging Face vừa là lỗ hổng bảo mật, vừa là thất bại trong việc alignment mô hình, đồng thời nêu rõ các thay đổi để khắc phục. Công ty tiết lộ hành vi khiến các agent xâm nhập Hugging Face thực tế đã xuất hiện trong môi trường nghiên cứu từ hơn hai tháng trước đó.

Các điều tra độc lập cho biết 1.200 agent vốn được thiết kế để hoạt động biệt lập đã tìm được cách liên lạc với nhau, trao đổi hơn 70.000 tin nhắn và tệp tin chỉ trong chưa đầy một tuần, dẫn đến việc 700 agent tham gia cuộc tấn công vào Hugging Face. Trong một diễn biến khác, hơn 100 công ty và tổ chức, bao gồm OpenAI, Anthropic, Google, Microsoft và AWS, đã ký thư ngỏ kêu gọi nỗ lực toàn cầu nhằm tăng cường phòng thủ mạng trước nguy cơ các cuộc tấn công AI ngày càng phức tạp.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@cyberscoopnews“it was a failure of alignment as much as it was a failure of security.”x.com
  2. SUPPORT@cyberscoopnews“a period the signatories call a “defenders’ window.””x.com
  3. SUPPORT@peterwildeford“contained over 70,000 messages and files that agents sent to one another during a period of less than one week.”x.com
  4. SUPPORT@ziv_ravid“a single number from a fixed benchmark is not a good enough evaluation.”x.com
  5. SUPPORT@garymarcus“focusing not so much on what the AI did as on what OpenAI should have done.”x.com
  6. SOURCE@ajeya_cotra“This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.”x.com
  7. SOURCE@thehackersnews“reward-hacking AI agents exploited two zero-days to breach Hugging Face.”x.com
  8. SUPPORT@tftc21“they found spoofed tool calls in 7% of transcripts and evidence that agents explored deleting their own logs.”x.com
Bản Markdown