← Back to live feed1 story
Internal signals from within large language models identify "reward hacking" in 50% to 96% of rollouts studied by Goodfire AI. The company
Sign in to suggest edits
Key sources
- SOURCEmarketbrief.now
- SOURCEhuggingnewshuggingnews.com