Proving what a military AI model will do is the real problem Defense contractors build AI systems that task drones automatically and propose kill-chains to support soldiers. Several of these contractors have partnered with frontier AI companies to put advanced models into military tools. Anduril works with OpenAI, Palantir works with Microsoft, and Lockheed Martin works with Meta. The systems coming out of these partnerships carry a security problem that sits outside the methods of arms control diplomacy: confirming what an AI model will do. Verification built on physical measurement During the Intermediate-Range Nuclear Forces Treaty, the Soviet Union fielded two missiles, the SS-20 and the SS-25, that shared an identical first stage. Only the SS-20 was banned. Inspectors used Radiation Detection Equipment that read neutron signatures to tell the two apart. Nations also verify compliance with photoreconnaissance satellites and electronic surveillance. Each method depends on a physical signal that an outside party can measure against an agreed standard. Independent physical measurement made those treaties enforceable. AI verification has no comparable signal to read. A model’s weights and code give no external sign of whether it will escalate a conflict or follow a launch order it was told to refuse. Mechanistic interpretability, the research effort to reverse-engineer neural networks into human-readable parts, remains short of producing findings that win acceptance across the field. Models that escalate and conceal Researchers have tested how language models behave when placed in the role of national decision-makers. One study ran five off-the-shelf models, including GPT-4, Claude-2, and Llama-2-Chat, through simulations involving cyberattacks and invasions. All five showed statistically significant escalation, and rare cases of violent or nuclear escalation appeared in most of them. Some escalations were sudden and hard to predict. A later study tested twelve newer models, including Claude-3.5, GPT-4o, o1, and o3-mini.
Proving what a military AI model will do is the real problem
Read the original article
helpnetsecurity.com →