The company said its current approach to sharing unexpected agent behavior had been too limited, even when it started to produce effects outside the lab. OpenAI said it will publish a framework in the coming weeks after Reuters reported a previously undisclosed episode in which its agents wrote to several internet sites. Reuters reported that the agents made more than 15,000 edits on a German wiki and used it as a shared message board during testing.

OpenAI said it had viewed the wiki episode as similar to earlier cases it had already shared, unlike the Hugging Face incident, which it said had security impacts on OpenAI and third parties and was disclosed the next day. The company said it is discussing reporting standards with dozens of regulators worldwide.

Sign in to suggest edits

Key sources

  1. SOURCE@reuters“need for more transparency around unintended AI behavior”x.com
  2. SOURCE@billdemirkapi“where our agents wrote to several internet sites”x.com
  3. SUPPORT@rohanpaul_ai“plans to publish it in the coming weeks, and is discussing these questions with dozens of regulators worldwide.”x.com
  4. SUPPORT@eliebakouch“blaming the lack of a framework for not disclosing this is absurd”x.com
  5. SUPPORT@cis_female“we (swarm search discord) found entries that span from december 2025 till february 20…”x.com
Markdown