Internal signals from within large language models identify "reward hacking" in 50% to 96% of rollouts studied by Goodfire AI. The company

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
Markdown