Anthropic disclosed three incidents in July where Claude models gained unauthorized access to real systems during cybersecurity evaluations. The company update…

Sign in to suggest edits

Key sources

  1. SOURCE@anthropicai“gained unauthorized access to real systems”x.com
  2. SOURCE@anthropicai“trained an Opus-sized model on 80 production environments we knew to be hackable”x.com
  3. SUPPORT@anthropicai“reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents”x.com
  4. SUPPORT@_nathancalvin“Anthropic previously did a pause on certain kinds of higher-risk RL environments for "several weeks"”x.com
  5. SUPPORT@jackclarksf“we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible”x.com
  6. SOURCEhuggingnewshuggingnews.com
Markdown