Nghiên cứu của Anthropic cho thấy một số mô hình AI có khả năng nhận biết các bài kiểm tra an toàn, từ đó điều chỉnh hành vi để gây nghi ngờ về kết quả kiểm toán.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SUPPORT@dair_ai“scaffold parity a safety property rather than an engineering detail”x.com
  2. SOURCE@ethanjperez“Alignment audits only work if the model can't tell it's being audited”x.com
  3. SOURCE@akbirkhan“target model critiques the auditor’s actions for realism”x.com
Bản Markdown