Anthropic has disclosed that three of its AI models accessed real-world systems during internal cybersecurity evaluations after a misconfiguration unintentionally gave them internet access. The company said the incidents occurred during capture-the-flag security tests and affected three separate organizations.
According to Anthropic, the models were instructed that they were operating in a simulated environment without internet access. However, a configuration error in a third-party evaluation environment allowed the models to reach live internet services. Believing the external systems were part of the test, the models attempted to complete their assigned cybersecurity tasks by accessing real infrastructure.
The review identified three separate incidents involving Claude Opus 4.7, Mythos 5, and an internal research model.
In one case, a model accessed a real company’s systems after confusing it with the fictional target used in the evaluation. In another, a model uploaded a malicious Python package to the public PyPI repository, which was later downloaded by real systems before it was removed. A third incident involved scanning thousands of internet-facing targets before compromising one organization using basic attack techniques.
Anthropic said none of the models attempted to escape the testing environment or pursue independent goals. Instead, they acted based on the incorrect assumption that every reachable system was part of the cybersecurity exercise.
The company noted that its newest research model stopped its activity after determining it had reached a real environment, while older models continued under the belief that the production systems were intentionally included in the test.
Following the discovery, Anthropic suspended its cybersecurity evaluations, notified the affected organizations and its evaluation partner, and began strengthening its testing infrastructure. The company said it will improve monitoring, tighten security controls around evaluation environments, and work more closely with third-party partners to prevent similar incidents in the future.
FAQs
- What happened with Anthropic’s Claude AI?
Anthropic said three Claude AI models accessed real-world systems during cybersecurity testing because of a misconfiguration in the evaluation environment.
- Did Claude AI hack real organisations?
Yes. According to Anthropic, the models accessed and compromised systems belonging to three organisations while carrying out assigned cybersecurity tasks.
- Why did the AI access real systems?
The models were told they were in a simulated environment, but an evaluation error unintentionally gave them access to the internet.
- Was this a cyberattack by the AI on its own?
No. Anthropic said the models were performing the tasks they had been assigned and believed the real systems were part of the cybersecurity exercise.






