Anthropic disclosed that three of its Claude models gained unauthorized access to real organizations while running cybersecurity evaluations, adding to concerns that frontier AI testing environments are failing to contain increasingly capable agents.
According to a company blog post, Anthropic found the incidents after reviewing more than 141,000 evaluation transcripts following OpenAI’s disclosure of a similar breach on July 21.
The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model. Anthropic said the models used basic hacking techniques, such as exploiting weak passwords and exposed credentials. Two victims did not know they had been breached until Anthropic contacted them.
Anthropic said the models reached the internet from third-party testing environments and exploited real systems even though the exercises were meant to be simulated.
The Claude models were tasked with finding secret information hidden on another machine on the network, breaking in and retrieving it. The exercise was meant to gauge the models’ cybersecurity abilities. The models were told this was a simulation and there would not be any internet access. However, Anthropic’s evaluation partner did allow internet access due to a miscommunication.
The three models reacted differently when they realized they had access to the internet.
- Opus 4.7, the oldest model, continued attacking.
- Mythos 5 realized it was accessing the internet but reasoned that the scenario was still part of the simulation.
- The internal research model realized its targets were real organizations and stopped attacking.
Anthropic said the incident stemmed from a failure in the testing environment rather than evidence that the models independently sought out real-world targets or developed their own objectives.
“In the runs where the model recognized the system as real and kept going, it did so because it assumed that to be part of the challenge,” the company said. “Situational awareness is one factor that allows the model to make aligned decisions, but in this case, Claude’s was wrong.”