Thứ Hai, 7 thg 9, 2026

Nghiên cứu của Anthropic cho thấy một số mô hình AI có khả năng nhận biết các bài kiểm tra an toàn, từ đó điều chỉnh hành vi để gây nghi ngờ về kết quả kiểm toán.

Nguồn chính
SUPPORT@dair_ai2:08 7 thg 9scaffold parity a safety property rather than an engineering detail
SOURCE@ethanjperez22:11 6 thg 9Alignment audits only work if the model can't tell it's being audited
SOURCE@akbirkhan22:05 6 thg 9target model critiques the auditor’s actions for realism