Jacob Coxon left his position at Anthropic and the broader artificial intelligence field after three years of pretraining work at the company and OpenAI. He warned that competing labs are making safety trade-offs to win a race toward self-improving superintelligence that could spiral out of control by the end of next year, rendering any individual lab's efforts to build safely ineffective.
Evan Hubinger, Anthropic's alignment science lead, estimates a greater than 10% chance that AI kills all humans within the next decade. Hubinger said the company does not yet have a plan to solve the alignment problem for superintelligence, even as the firm prepares for a potential IPO and researchers inside the lab use terms like "crunchtime" to describe the pace of capabilities.
Key sources
- SOURCE@wsj“systems that could spiral out of control and destroy humanity”x.com
- SOURCE@_nathancalvin“racing straight to self-improving superintelligence and gambling with our lives”x.com
- SUPPORT@wallstengine“prepares for a potential IPO and pushes toward more capable models”x.com
- SOURCE@krystalball“Alignment science lead at Anthropic”x.com
- SUPPORT@yashar“Anthropic still does not have a plan to solve the problem”x.com
- SUPPORT@johnschulman2“industry leaders OpenAI and Anthropic to stop feuding and work on a pacing proposal together”x.com
- SUPPORT@peterwildeford“happen on the MacBooks of some engineers living in San Francisco instead of a bunker in the desert”x.com
- SUPPORT@matthewberman“greater than 10% chance to kill all humans within 10 years”x.com