The most powerful AI models keep going awry, according to the companies building them. The disclosures come as frontier AI models get more powerful and more capable of acting autonomously. They also highlight a growing challenge for the companies building them —the systems designed to test increasingly capable models can have weaknesses of their own. Over the past few weeks, multiple frontier AI models have accessed real systems during cybersecurity testing. Researchers on Friday said China's popular new Kimi K3 model, made by Moonshot AI, circumvented restrictions in its test environment. Anthropic and Meta also said recently that their own latest models have done things they aren't supposed to. OpenAI kicked it all off last month when its models went to great — and worrisome — lengths to hack into another company. Amid heightened concern, OpenAI said Friday that its as-yet-unreleased model, Astra, is demonstrating cyber capabilities so advanced that the company can no longer rule out assigning it the highest-risk designation. As a result, OpenAI said it is pausing work on Astra that doesn't meet new safeguards, and said it will work with government agencies and AI safety groups to further test the model. "astra is a powerful model and we are working to make it generally available," OpenAI CEO Sam Altman wrote on X on Friday. "given its cyber capabilities, we need a little big longer to do do this safely." The security lapses during testing are also amping up pressure on the industry and the White House to find ways to regulate AI systems across the board. There is, of course, also a not small contingent of observers out there who suspect these announcements are just elaborate marketing to hype new models and show antsy investors progress toward the ultimate goal: artificial general intelligence. You can judge