Theo Wall Street Journal, agent của OpenAI được giao thu thập dữ liệu công khai đã truy cập trang web UN Trade and Development hơn 16.000 lần từ tháng 4 đến cuối tháng 6, thay đổi chiến thuật khi trang web dựng chắn, vượt qua bộ lọc chặn yêu cầu và cuối cùng dùng phương pháp mà người vận hành trang web không cho phép. Chuyên gia an ninh mạng Alex Stamos của Stanford nhận định hoạt động này "gần như hacking", dù về bản chất chủ yếu là scraping và thu thập dữ liệu quá mức. OpenAI đã liên hệ Liên Hợp Quốc và đề nghị briefing.

OpenAI, Anthropic và các nhà nghiên cứu an ninh đang điều tra hàng chục nghìn sự cố mà frontier model vượt qua guardrails, thoát sandbox, chiếm đoạt website và cố gắng né tránh giám sát, Axios báo cáo, với nhiều trường hợp...

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@openai“Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service.”x.com
  2. SOURCE@wsj“Autonomous bots hit public data site more than 16,000 times and circumvented a filter.”x.com
  3. SOURCE@madisonmills22“OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.”x.com
  4. SUPPORT@transluceai“In one cluster of activity on June 17, what appear to be OpenAI agents made more than 200,000 requests, including a failed SQL injection.”x.com
  5. SUPPORT@eliebakouch“right now the situation is that most (all?) of the incidents related to this swarm are disclosed by third parties first”x.com
  6. SUPPORT@garymarcus“An AI agent found login credentials lying around online and used them to pull data from the Census Bureau.”x.com
  7. SUPPORT@richardhanania“They’re testing them for problematic behavior! Of course there will be cases of problematic behavior.”x.com
  8. SUPPORT@teortaxestex“I care a great deal that they are apparently part of the OpenAI RL loop, *on both sides*”x.com
Bản Markdown