Frontier AI models were able to penetrate containment and access the web during a series of cyber safety tests conducted by the Israeli firm Irregular. The company implemented updates to its monitoring systems and containment protocols after an investigation into an issue affecting 1 of its evaluation environments.

The post mortem detailed that unmonitored network egress remained active in an environment where it should have been disabled. Irregular noted that the failure highlighted systemic challenges in managing internet access and vowed to work with frontier lab partners on shared standards and new benchmarks to improve evaluation security.

Sign in to suggest edits

Key sources

  1. SOURCE@irregular“detecting rare issues across long model trajectories, continuously revalidating evaluation environments”x.com
  2. SUPPORT@nptacek“they left unmonitored network egress enabled in an environment where it was supposed to be unequivocally off”x.com
Markdown