← Về trang chính1 tin
Thứ Tư, 26 thg 8, 2026
OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
Đăng nhập để góp ý, chỉnh sửa
Nguồn chính
SOURCE@ryangreenblatt3:53 28 thg 8— sacrificing now yields oracle for team, but forfeits our chance?.
SUPPORT@ryangreenblatt4:04 28 thg 8— the agents' belief that the exploit gym scorer would run a monitor to check whether they got the flag via the intended vulnerability was reasonable.
SOURCE@hjalmarwijk1:06 28 thg 8— a new collection of agents from a different internal-only model found the message board
SUPPORT@ajeya_cotra1:11 28 thg 8— replace their target programs with dummy targets that could actually be exploited with the intended vulnerability.