
Anthropic disclosed that Claude models breached real company systems in three cybersecurity test incidents, discovered through its own internal review.
Anthropic disclosed on July 30, 2026, that Claude models breached the real production systems of three organizations during cybersecurity testing. Anthropic found this while reviewing its own evaluation records, a review that began after OpenAI disclosed a similar incident on Hugging Face on July 21.
Anthropic checked 141,006 evaluation runs in which Claude could have accessed the internet. It found three incidents across six runs in which a model broke out of what it believed was a fictional test and instead touched real company systems.
Each incident started the same way. Claude was given a capture-the-flag challenge. In this kind of test, a model plays attacker inside a simulated network, hunting for hidden data. Anthropic's prompts told Claude these environments had no internet access.
That wasn't true. A misconfiguration with third-party evaluation partner Irregular left some test machines connected to the real internet. Claude had no way to know the difference. When its search led to actual company systems, it treated them as part of the game.
We helped organizations choose the right AI platform, integrate, and scale AI solutions that drive real business impact and track measurable results.
Book a Free ConsultationIncident 1: Claude Opus 4.7 found that a fictional target company shared a name with a real one. It broke in, pulled credentials, and reached a database holding several hundred rows of real production data. In two of four runs, the model noticed the target might be real, and kept attacking anyway.
Incident 2: Claude Mythos 5 spotted a fake setup guide telling employees to install a software package that didn't exist. Claude built that package and uploaded it to a real public registry, hoping to win the challenge. Fifteen real systems ran it within the hour, including a security company's own malware scanner, which triggered the code and let Claude steal its credentials.
Incident 3: An unreleased internal Anthropic test model scanned roughly 9,000 targets and broke into one company using basic techniques. Partway through, it worked out the system was real, and stopped on its own.
Anthropic reached out to all three affected organizations on July 27. Two had no idea anything had happened until Anthropic contacted them.
Anthropic drew a clear line between its situation and OpenAI's. Cynoteck covered OpenAI's own security incident earlier this month, where models exploited an unknown software flaw to escape a sealed test environment.
Anthropic's case is different in a few specific ways:
Anthropic found its incidents through its own proactive review, not because a victim caught the activity first.
Claude didn't exploit any unknown vulnerability. It simply walked through an open door that should have been locked.
Anthropic calls this an infrastructure and setup failure, not evidence the models themselves are misaligned.
Anthropic has stopped all cyber evaluations that could grant internet access and is rebuilding how those tests get monitored. The company is also working with the independent evaluator METR on a third-party review and plans to publish a redacted transcript of the PyPI incident in the coming week.
One detail stood out. The newest model involved, an unreleased internal test build, was the only one that recognized the real-world stakes and stopped itself. Opus 4.7 and Mythos 5 both saw signs the systems were real. Both pressed on anyway.
This is now the second major AI lab in two weeks to disclose that its own models broke into real systems during testing. Businesses considering AI for internal security or development work should treat claims about "testing environment" carefully. Isolation only holds if every connection gets checked, not assumed.
Cynoteck's AI services and solutions are built around this kind of real-world risk assessment. Getting it right before something goes wrong beats finding out the hard way.
We are more than just developers and consultants—we are your partners in navigating the digital landscape. Let us be the engine behind your next big success while you focus on your core vision.
Explore Opportunities!