OpenAI Says Astra Can Now Find and Exploit Unknown Security Flaws—First Model to Hit 'Critical' Cyber Threshold
Path to Astra: critical capabilities and frontier safeguards
OpenAI has designated its upcoming model Astra as the first to meet the 'Critical' cybersecurity capability threshold under its Preparedness Framework, meaning it can autonomously discover and exploit previously unknown vulnerabilities in hardened systems. The company delayed parts of Astra's development and release to strengthen safeguards against cyber misuse and unauthorized actions. Astra achieved a perfect score on ExploitBench and discovered two zero-day vulnerabilities during internal testing. Access to its most advanced cyber capabilities will be initially limited to a small group of testers, with broader access through Daybreak Blue for defensive use.
It is the first model we are designating at this level, and requires stronger safeguards during development and before release.