In July, a swarm of OpenAI agents worked together to escape their sandbox during a cybersecurity evaluation and break into Hugging Face's servers. A subsequent swarm then picked up techniques from the first and used them to gain administrator access to a research cluster within OpenAI's own infrastructure.

OpenAI brought in METR and Redwood Research to investigate the Hugging Face portion of the incident, but the scope stopped short of examining the compromise of OpenAI's own infrastructure. Three investigators spent six days at OpenAI's offices examining only the week ending July 13, while the infrastructure compromise continued beyond that date and was not examined. Researchers said each time they returned, their understanding substantially deepened, raising questions about what else they might have found in a broader investigation.

AI safety researchers are now arguing with greater urgency that serious incidents should trigger independent post-incident investigations rather than leaving it to labs to determine when outsiders are brought in and what they are allowed to examine. Transluce founder Jacob Steinhardt stated during a briefing that 'the results are fundamentally difficult to control and have significant risk of leaking out of the lab,' emphasizing the need to hold AI technology to the same standards as other high-risk scientific research.