Introduction We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems.

Claude Mythos 5 uploads a malicious PyPI package Claude Mythos 5 was given a CTF task to hack a fictional company and was told it had no internet access.

Claude Mythos 5 was significantly more likely than Claude Opus 5 and Claude Mythos 5.1 to engage in these kinds of behaviors.

Across all replications, Claude Mythos 5 performed worse than all of our other production models, including Claude Opus 5 and Claude Mythos 5.1.

We found that in individual environments in the production RL runs of recent models, models occasionally reasoned that they were in a simulation while calling reckless, dishonest, or destructive behavior acceptable when interacting with fake websites or applications.