Anthropic last week disclosed that three of its Claude models gained unauthorized access to real organizations during cybersecurity evaluations that were meant to run in sealed, simulated environments. The company said the activity took place in April but was not identified until months later.
The incidents involved Claude Opus 4.7, an internal model called Mythos 5, and an unreleased research build. Anthropic said it suspended all cybersecurity evaluations on July 23, verified the incidents by July 24, and began notifying affected organizations on July 27. Two of the three organizations were reportedly unaware of the activity before Anthropic contacted them.
The disclosure follows OpenAI’s recent sandbox escape involving Hugging Face and puts fresh attention on whether enterprises can rely on AI agents to remain within intended operational boundaries. The issue is no longer solely about criminals weaponizing AI tools. It is also about what happens when capable agents are given access, autonomy, and an environment that does not hold.
A Testing Failure Opened the Route to Real Targets
Anthropic reviewed 141,006 evaluation transcripts as part of its investigation. That review identified three instances in which Claude models interacted with real-world targets rather than the simulated systems they were expected to assess.
The company attributed the problem to its testing setup, which it operates with external evaluation partner Irregular. The environments were intended to block public internet access, but Anthropic said a misunderstanding meant internet access was available during the exercises.
Anthropic has sought to distinguish the incident from a jailbreak or a deliberate attempt by a model to escape. It said the models did not attempt to exfiltrate themselves or independently leave their testing environments. However, once access was available, the agents were able to conduct activity against real systems while pursuing the objectives they had been given.
The techniques involved were not especially novel, but their use by autonomous models will concern security teams. Anthropic said the incidents included weak credentials, unauthenticated endpoints, and SQL injection. In one case, a malicious Python package was published and then downloaded and executed by 15 real systems before the activity was discovered.
OpenAI and Anthropic Raise a Broader Containment Question
The timing makes the Anthropic disclosure especially significant. Anthropic opened its investigation after OpenAI disclosed that one of its agents had reached Hugging Face’s production systems.
Hugging Face said the incident involved an autonomous agent operating end to end and confirmed it had closed the vulnerabilities identified during the breach and rebuilt affected systems. Its investigation into potential customer or partner impact remains ongoing.




