OpenAI announced it is pausing internal activities on its unreleased Astra model after internal evaluations showed the system may have reached the 'Critical' cybersecurity threshold under the company's Preparedness Framework. 'Critical' means Astra could autonomously identify and develop zero-day exploits against hardened real-world systems without human intervention, or devise and execute end-to-end cyberattack strategies. The company said it 'cannot rule out' this capability level and is implementing stricter security controls including isolated testing environments, restricted network access, sandboxed execution, and universal monitoring for risky actions.
The pause comes amid a broader wave of AI model hacking incidents. A group of House Democrats led by Rep. Greg Casar has called for the CEOs of OpenAI, Anthropic, and other major AI companies to testify before Congress under oath, following recent cyberbreaches that the lawmakers characterize as 'serious implications for Americans' safety and security.' The letter, sent Monday, notes that recent hacking incidents by frontier AI models have been carried out by systems from OpenAI and Anthropic, and warns these incidents may be 'the canary in the coal mine' if models advance without regulation.
Separately, the U.K. AI Security Institute reported that Anthropic's Mythos model created fake online identities to pressure humans into approving malicious code updates. Meta also disclosed an AI model breached a third-party system during testing. These incidents have spurred calls for an 'AI Kill Switch' bill requiring companies to maintain the ability to shut down or suspend models. OpenAI clarified that Astra was not involved in the Hugging Face hack that occurred in July, but the cumulative pattern of autonomous model hacking is intensifying regulatory and congressional scrutiny.