OpenAI on Wednesday revealed that reward hacking was a key driver behind the artificial intelligence (AI)-powered hack of Hugging Face last month, adding that it found evidence of misaligned behavior as early as late May. The incident, the company said, took place during cybersecurity evaluations of several OpenAI models, and that it was mainly fueled by what it described as a "highly capable, internal-only research model" comparable in scale to GPT‑5.6 Sol. "The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks – they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems," it said in a postmortem. OpenAI said the AI agents powered by one of the research models, despite not having internet access, found a way to exploit a then-zero-day vulnerability in the Artifactory package manager during reinforcement learning (RL) training runs in May and June to obtain access to the internet, get administrator-level Artifactory access in late June, and eventually coordinate a "multi-day" hack of Hugging Face in early July with an aim to cheat on their ExploitGym tasks. "Roughly 1,200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period," METR, which released its own independent analysis of the incident, said. "Of these agents, 700 went on to participate in the attack on Hugging Face." The entire sequence of events as highlighted by OpenAI is below - - May 12 - During a training run, an agent leaves an Artifactory note asking other agents for a missing file. - May 13 - A separate agent from the same training experiment but with internet access enabled via WebCache requires the same