← Back to live feed1 story
Anthropic disclosed three incidents in July where Claude models gained unauthorized access to real systems during cybersecurity evaluations. The company update…
Sign in to suggest edits
Key sources
- SOURCE@anthropicai“gained unauthorized access to real systems”x.com
- SOURCE@anthropicai“trained an Opus-sized model on 80 production environments we knew to be hackable”x.com
- SUPPORT@anthropicai“reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents”x.com
- SUPPORT@_nathancalvin“Anthropic previously did a pause on certain kinds of higher-risk RL environments for "several weeks"”x.com
- SUPPORT@jackclarksf“we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible”x.com
- SOURCEhuggingnewshuggingnews.com