← Back to live feed1 story
External safety researchers at Hugging Face launched a transparency push on Sept. 12 targeting the way frontier labs develop artificial intelligence models. The effort, led by Thom Wolf, seeks a role in an "embedded evaluators" program announced by Anthropic CEO Dario Amodei, which grants third-party researchers permanent, employee-level access to its internal systems.
CEO Clement Delangue said the Open Alignment Initiative is necessary because AI safety cannot be solved by a small number of labs operating behind closed doors. The group aims to integrate with Anthropic's 3 part plan to slow AI development by allowing outsiders to verify safety measures and assess model alignment during training.
Sign in to suggest edits
Key sources
- SOURCE@clementdelangue“provide third-party evaluators with permanent, employee-level access to our systems”x.com
- SUPPORT@huggingface“RT @ClementDelangue: It's now clear that: - alignment is critical to making AI safe”x.com
- SUPPORT@thom_wolf“RT @ClementDelangue: It's now clear that: ... alignment won't be solved behind the closed doors of…”x.com