AI war games almost always escalate to nuclear strikes, simulation shows A new study reveals that AI decision-making during conflicts is naturally prone to escalation. Get the world’s most fascinating discoveries delivered straight to your inbox. You are now subscribed Your newsletter sign-up was successful Want to add more newsletters? Join the club Get full access to premium articles, exclusive features and a growing list of member rewards. Defense and intelligence agencies are increasingly relying on artificial intelligence (AI) systems to augment their capabilities, including for pattern recognition in intelligence gathering and scenario planning for contingency operations. Yet one of the core issues of AI and large language models is that we have never truly understood the logic underpinning them, scientists say. These systems have been compared to a black box that provides answers without showing the reasoning to support the outcomes. To understand the logic of AI systems, Kenneth Payne, a professor of strategy at King's College London, designed a series of war gaming simulations between two competing AIs and found that in nearly every scenario, nuclear escalation was unavoidable. He published his findings, which have not been peer-reviewed, Feb. 16 in the arXiv preprint database. The experiment used a series of two-way tournaments of the Khan Game, in which Claude Sonnet 4, GPT-5.2 and Gemini 3 Flash competed in a series of simulated nuclear crises. The Khan Game is an AI-vs-AI strategic escalation simulation between two nuclear powers, with state profiles loosely based on the Cold War. One is technologically superior but militarily weaker, while the other is militarily stronger but adopts a risk-tolerant leadership style. Some of the simulations included allied nations, with one scenario deliberately testing whether an alliance leadership could be maintained during the conflict. Each turn, the AIs simultaneously signaled their intentions before they