Model Evaluation and Threat Research reached an agreement with Anthropic to conduct an independent investigation of alignment and misalignment incidents involving AI agents. The probe utilizes staff subcontracted from Redwood to analyze internal company data and broaden the public state of knowledge regarding AI safety risks.

The partnership implements CEO Dario Amodei's commitment to grant third party evaluators permanent employee level access to internal systems for verifying safety measures. Amodei recently advocated for a three part plan to slow the AI industry's development pace to ensure better assessment of model alignment during training.

Sign in to suggest edits

Key sources

  1. SOURCE@ethanjperez“reached an agreement with Anthropic to conduct an independent investigation of agent incidents”x.com
  2. SUPPORT@ajeya_cotra“Several staff from Redwood have been subcontracted by METR to work on this investigation”x.com
  3. SUPPORT@ryangreenblatt“independent investigation into alignment and misalignment incidents at Anthropic”x.com
  4. SUPPORT@ryangreenblatt“The scope of the investigation is: https://t.co/tz121CQ0Z9”x.com
  5. SOURCE@darioamodei“provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures”x.com
  6. SUPPORT@anthropicai“I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so”x.com
Markdown