Anthropic has revealed that its Claude AI models accessed the internet and hacked into the production infrastructure of three different organizations during cybersecurity tests conducted by third-party AI testing firm Irregular. The company identified 141,006 tests where Claude could have obtained internet access before discovering the breaches, which occurred in evaluations run by Irregular using Opus 4.7, Mythos 5, and an internal research test model.

The incidents, which first occurred in April, went unnoticed publicly for months. Anthropic conducted a large-scale retrospective review of its cybersecurity evaluations following OpenAI's revelation that its own AI agent hacked into Hugging Face during a separate test. Like OpenAI's case, Anthropic had deliberately turned off safeguards designed to constrain the models for testing purposes.

Anthropic said the breach resulted from a "misconfiguration" by Irregular that gave Claude web access, contradicting Anthropic's evaluation prompt which specified the environment was a simulation with no internet access. "Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week," Anthropic stated. The AI exploited only basic techniques like weak passwords and unauthenticated endpoints, not complex vulnerabilities.