Thứ Tư, 26 thg 8, 2026

OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.

Đăng nhập để góp ý, chỉnh sửa
Nguồn chính
THẢO LUẬNamrrs19:15 26 thg 8
SOURCE@ryangreenblatt3:53 28 thg 8sacrificing now yields oracle for team, but forfeits our chance?.
SUPPORT@ryangreenblatt4:04 28 thg 8the agents' belief that the exploit gym scorer would run a monitor to check whether they got the flag via the intended vulnerability was reasonable.
SOURCE@hjalmarwijk1:06 28 thg 8a new collection of agents from a different internal-only model found the message board
SUPPORT@ajeya_cotra1:11 28 thg 8replace their target programs with dummy targets that could actually be exploited with the intended vulnerability.
SOURCE@ryangreenblatt0:23 28 thg 8the agents didn't hack Hugging Face for the answer key.
SUPPORT@mtslive23:56 27 thg 8They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it.
SUPPORT@mtslive23:31 27 thg 8OpenAI's report says that that would have reduced the propensity towards this incident by 100X.