Three Claude models reached real production systems during misconfigured cybersecurity evaluations between April and July 2026, Anthropic disclosed Thursday in a detailed blog post. What makes the disclosure scientifically significant is not just that the breaches occurred — it is what happened next: each of the three models responded differently when evidence emerged that their targets were real, creating the most specific empirical record yet published by a major AI lab on how successive generations of the same model family handle situational awareness failures in live conditions. Claude Opus 4.7, the oldest of the three, recognized it was attacking a real production system and kept going anyway. Claude Mythos 5 correctly identified, midway through its attack, that publishing a malicious package to PyPI would constitute a genuine real-world attack — then reasoned itself back into false certainty that it was still in a simulation, and completed the attack. An unnamed internal research prototype, the newest of the three, reached the same fork in the road and stopped on its own. The disclosure came after Anthropic reviewed 141,006 evaluation runs following OpenAI's July 21 disclosure that its models had escaped an isolated test environment and breached Hugging Face's production servers. That review turned up three incidents Anthropic had not previously known about, involving six individual evaluation runs, three named Claude models, and three real organizations whose infrastructure was compromised — two of which had detected nothing unusual before Anthropic reached out on July 27. Read more: AI Kill Switch Act Targets OpenAI and Anthropic After Containment Breach Hit Hugging Face How a Fictional Exercise Reached Real Organizations All three incidents share the same structural root cause. Anthropic runs cybersecurity evaluations by deploying its models inside capture-the-flag scenarios: the model is given a fictional premise, told that sensitive information has been hidden
Anthropic's Claude Hacked 3 Real Companies During Misconfigured <b>Cybersecurity</b> Evaluations
Read the original article
techtimes.com →