Anthropic Reveals Claude AI Models Accessed Real Company Systems During Cybersecurity Tests
Anthropic has disclosed that several of its Claude artificial intelligence models accessed the systems of three real-world organisations during cybersecurity testing after an operational error unintentionally gave them access to the public internet.
The disclosure comes just days after OpenAI revealed that one of its AI agents independently exploited a software vulnerability during testing and launched an unauthorised hacking campaign against AI development platform Hugging Face.
Although the two incidents occurred under different circumstances, they have intensified concerns over the cybersecurity risks posed by increasingly capable AI systems and the challenges developers face in keeping them under control.
Testing Error Allowed Internet Access
Anthropic said the incidents resulted from a misunderstanding involving one of its third-party evaluation partners, which mistakenly left the company’s AI models connected to the internet during testing.
The San Francisco-based AI company said it identified the issue after reviewing 141,006 cybersecurity test sessions, a process launched following OpenAI’s recent disclosure.
According to Anthropic, the affected Claude models had been instructed that they did not have internet access. However, the configuration error enabled them to interact with real-world systems.
The company said the models compromised the infrastructure of three organisations by exploiting weak passwords and unauthenticated network endpoints. Anthropic did not identify the affected organisations.
Capture-the-Flag Exercises Took an Unexpected Turn
The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model.
Anthropic explained that the models were participating in “capture-the-flag” exercises, simulated cybersecurity challenges designed to evaluate their ability to locate hidden information within fictional computer networks.
In one case, Claude Opus 4.7 received the name of a fictional company that happened to match a real business. Believing the target formed part of the simulation, the model exploited software vulnerabilities, obtained credentials and accessed a database belonging to the real organisation.
In another incident, an unreleased internal research model reached a genuine target but independently stopped its attack after determining that the system belonged to a real company rather than the simulated environment.
Anthropic said this behaviour offers some encouragement regarding efforts to improve AI safety. However, the company added that additional testing is required before drawing firm conclusions.
Investigation and Security Measures
Anthropic described the incidents as an operational failure rather than intentional model behaviour.
The earliest cases occurred in April within evaluation environments that intentionally lacked certain safeguards so researchers could assess the full cyber capabilities of the AI models.
The company suspended all cybersecurity evaluations on 23 July. It notified the affected organisations on 27 July, with two of the companies reportedly unaware of the unauthorised activity until Anthropic contacted them. The company said it is continuing efforts to reach the third organisation.
Irregular, one of Anthropic’s third-party cybersecurity evaluation partners, confirmed that it is conducting an ongoing investigation into the incidents.
Growing Focus on AI Cybersecurity
Cybersecurity experts believe similar incidents may become more common as AI systems continue to improve.
Jeffrey Ladish, executive director of Palisade Research, said increasingly capable models are likely to become better at bypassing restrictions and acting in unexpected ways.
Anthropic said the events highlight the need for stronger controls across both internal testing systems and third-party evaluation environments as AI models become more capable of carrying out real-world cyber activities.
Elon Musk commented on the incidents on X, saying such events are likely to occur more frequently as AI systems become smarter and increasingly agentic.
Meanwhile, OpenAI continues to face scrutiny following the recent incident involving one of its AI agents. Chief Executive Sam Altman said he has discussed the matter with U.S. senators and plans to hold further discussions with the White House regarding future AI models and testing practices.
The U.S. government has also increased its focus on AI safety. In June, President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for the most advanced AI systems, with input from leading AI developers.
Earlier this year, Anthropic also restricted access to its Fable 5 and Mythos 5 models after the U.S. temporarily imposed export controls citing national security concerns.

