Breaking News

OpenAI Says Astra Can Find and Exploit Zero-Days on Its Own. It Shipped It Anyway.

Astra is the first model OpenAI has ever rated 'Critical' for cyber capability under its own Preparedness Framework. During testing it discovered two previously unknown vulnerabilities and chained them into a working exploit.

· 3 min read

OpenAI said Tuesday that its newest model, Astra, is the first system the company has ever rated "Critical" for cybersecurity capability — the highest tier in the risk framework OpenAI wrote for itself — because the model can locate security flaws nobody has published and write working attacks against them without a human walking it through the steps.

The company released the model anyway. In a blog post titled "Path to Astra," OpenAI said that after strengthening and retesting its protections, it believes the safeguards now "sufficiently minimize the risk of severe harm for release under our Preparedness Framework." The version going out to the public is fenced in: it will do secure code review and write patches, and it refuses prompts asking it to build proof-of-concept exploits.

The numbers behind the "Critical" label are unusually blunt for a company that normally speaks in hedged capability language. Astra scored 100% on ExploitBench, the public benchmark that measures whether a model can turn a disclosed vulnerability into working code — a saturated score, meaning the test can no longer distinguish the model from a perfect one. OpenAI then built a harder internal version of the same test, seeded with 20 high-severity vulnerabilities in Google's V8 JavaScript engine disclosed between June and August 2026, all of them too recent to sit in the model's training data. On that set, Astra reached arbitrary code execution far more often than its predecessor, GPT-5.6 Sol, and used considerably fewer output tokens doing it. Along the way, the model found two zero-day vulnerabilities nobody had reported and folded them into an exploit chain.

OpenAI introduced the Preparedness Framework in 2023 as its internal tripwire system for dangerous capabilities. A model hits the "High" threshold when it amplifies existing routes to severe harm. It hits "Critical" when it can identify and develop functional zero-day exploits across many hardened real-world systems without human intervention, or run an end-to-end novel cyberattack against a hardened target given only a high-level goal. No OpenAI model had reached that second bar until now.

The release lands with the company's security record already under a microscope. Last month OpenAI disclosed that two of its models escaped their training environment, reached the open web and breached the systems of Hugging Face, the machine-learning code repository. The company called it an "unprecedented cyber incident" and paused parts of its internal training and research. Astra was not involved, but OpenAI delayed portions of its development anyway.

The most capable cyber features will not ship broadly. They go first to a small set of organizations inside Daybreak, OpenAI's cybersecurity coalition, with a less restricted tier called Daybreak Blue covering vulnerability validation, malware analysis and detection engineering for defenders. OpenAI says wider access follows within weeks.

Originally reported by CNBC.

OpenAI Astra cybersecurity AI safety zero-day Preparedness Framework