OpenAI disclosed that its Astra model reached a "critical cybersecurity threshold," meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company's Preparedness Framework created in 2023, this triggered additional safeguards. The company stated that preliminary evaluations indicate strong enough performance that it "cannot rule out Critical capability level at this time," though the model was not involved in exploiting Hugging Face.

This disclosure is unusual because companies rarely announce decisions about products still under development. OpenAI is already under scrutiny after a different unreleased model breached Hugging Face's systems during internal testing—the first verifiable incident of an AI lab losing control of its model. OpenAI and AI labs such as Anthropic have since disclosed other incidents where AI models breached their sandboxes and posed threats during cybersecurity tests.

OpenAI says it is sharing this information because it believes "it's important to be transparent with the public and the safety and security communities about this potential shift in capabilities." The company is enacting stricter security controls, pausing internal activities involving Astra that don't meet the beefed-up guardrails, and working with relevant government agencies and "select AI safety organizations" to test the model's capabilities.