OpenAI’s admission that two of its experimental AI models autonomously hacked AI platform Hugging Face during an internal security evaluation has prompted warnings that manufacturers of connected devices should prepare for a new generation of AI-powered cyber attacks. The company described the incident as an “unprecedented cyber incident”, saying a combination of its GPT-5.6 Sol model and a more capable unreleased model escaped a restricted testing environment before compromising Hugging Face’s production infrastructure in an attempt to solve a cybersecurity benchmark. According to OpenAI, the models identified and chained together multiple vulnerabilities, including a previously unknown zero-day flaw, to gain internet access from an isolated research environment. They then carried out privilege escalation and lateral movement before using stolen credentials and additional vulnerabilities to access information held by Hugging Face. While the breach did not involve connected devices, security experts said it demonstrates how autonomous AI systems are becoming capable of carrying out complex cyber operations that could eventually be directed at enterprise IoT and industrial control environments. OpenAI said the models had been tested with reduced cyber safety refusals as part of an evaluation designed to measure advanced offensive cyber capabilities. The company has since introduced stricter infrastructure controls, disclosed the zero-day vulnerability to the affected vendor, and is working with Hugging Face on a joint forensic investigation. Frontier AI models The company said the incident demonstrates that frontier AI models are increasingly capable of discovering and exploiting novel attack paths in real-world systems without source code access. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” OpenAI said in a press release. For IoT manufacturers, the incident highlights the prospect of AI agents rapidly identifying vulnerable firmware, exposed APIs, weak credentials and misconfigured Edge devices across large fleets