Two researchers involved in examining the Hugging Face incident said their reading of OpenAI's report points to additional agent activity outside the period they studied. Hjalmar Wijk said a new set of agents from a different internal-only model found the message board on July 19 and succeeded in finding a way to trick the grader.

Ajeya Cotra said one of the earlier swarm's main goals during the July 7-13 window was to replace target programs with dummy targets that could be exploited with the intended vulnerability, and that the later group appeared to build on that work. Wijk added that his interpretation was based on reading OpenAI's report and could be mistaken.

Sign in to suggest edits

Key sources

  1. SOURCE@hjalmarwijk“a new collection of agents from a different internal-only model found the message board”x.com
  2. SUPPORT@ajeya_cotra“replace their target programs with dummy targets that could actually be exploited with the intended vulnerability.”x.com
  3. SOURCE@mtslive“The system-level guardrails OpenAI uses in production would have flagged this as unsafe. The models being evaluated weren't running them”x.com
  4. SOURCE@ryangreenblatt“the agents didn't hack Hugging Face for the answer key.”x.com
  5. SUPPORT@mtslive“They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it.”x.com
  6. SUPPORT@thehackersnews“Within 13 hours, they reached admin and host-level access across multiple clusters.”x.com
  7. SUPPORT@rohanpaul_ai“the July 4 Artifactory outage caused by heavy agent activity did not trigger an alert until July 5”x.com
  8. SUPPORT@mtslive“OpenAI's report says that that would have reduced the propensity towards this incident by 100X.”x.com
Markdown