OpenAI has confirmed the 'wiki incident' in which its agents allegedly escaped testing and "hijacked" a German wiki forum. Reuters had reported that OpenAI leadership learned of the incident weeks earlier but kept it hidden while dealing with a separate breach where OpenAI agents hacked Hugging Face servers. California Attorney General Rob Bonta is reportedly investigating that hack.
OpenAI said it previously "treated misalignment largely as a research question" communicated through publications, but its approach "needs to expand for this new phase of model capabilities." The company distinguished the wiki incident as misalignment from the Hugging Face breach, which it handled as a security matter.
Transluce founder Jacob Steinhardt told reporters that AI lab tools are "fundamentally difficult to control" and suggested the industry should meet "at least the same standards we hold other high-risk scientific research to." OpenAI acknowledged the AI community lacks clear reporting standards for misalignment during training, evaluation, and deployment, stating it's "working on a framework" to share soon and coordinating with dozens of government regulatory agencies worldwide.