AERIOXFLUX
Frontier Labs
Frontier Labs · openai

OpenAI Shipped the Model It Paused and Labeled It Critical

Astra scored 100% on ExploitBench, found two zero-days during testing and chained them, and became the first model OpenAI has classified as Critical for cybersecurity under its own Preparedness Framework. It shipped anyway.

Flux Desk·2026-09-07·5 min read

On August 9, Flux covered OpenAI suspending parts of Astra after internal cybersecurity testing flagged the model's offensive capability. It read at the time like the first case of a capability threshold producing an actual product hold rather than a policy memo.

On September 3, OpenAI shipped it. GPT-6 Astra launched as what the company calls its most intelligent and most aligned model, and simultaneously as the first model ever to cross the "Critical" cybersecurity threshold under OpenAI's own Preparedness Framework.

Both things are true, and the second one is the story.

What "Critical" means here

Under the Preparedness Framework, Critical is the top rung. For cybersecurity it means, in substance, that the model can independently discover and exploit previously unknown vulnerabilities in well-defended systems given appropriate tools and access. Not assist a researcher. Not summarize a CVE. Find the bug and build the chain.

OpenAI's evidence for the classification is its own testing. During evaluation, Astra found two zero-day vulnerabilities and used them in an exploit chain. OpenAI says it is disclosing those to the affected software maintainers.

The benchmark number is the blunt version. Astra saturates ExploitBench at 100%. Its predecessor, GPT-5.6 Sol, launched at 73.5% and climbed to 78.5% over its lifecycle. That is not an incremental generation. That is a benchmark being retired by a single release.

The rest of the scorecard is in the same register: reported launch figures of 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4, and 72.6% on OSWorld 2.0, alongside third-party tracking showing 98.5% on ARC-AGI-1, 100% on MRCR v2 at 256K–512K, and 92.7% on ScreenSpot Pro.

Shipping a Critical model

The interesting design decision is not the capability. It is how OpenAI split access around it.

The public model carries production safeguards. Astra will refuse the sharper cybersecurity requests — most visibly, it declines to write proof-of-concept exploits. The capability is present; the surface is narrowed by refusal behavior rather than by withholding the model.

Then there is OpenAI Daybreak, a program the company says will loosen those restrictions in the coming weeks for vetted defenders. Approved organizations get a less-restricted Astra for defensive work: vulnerability validation, malware analysis, and detection engineering. Daybreak participants are also first in the rollout order, ahead of ChatGPT Plus, Pro, Business and Enterprise, the API, and AWS.

If that structure sounds familiar, it should. It is the same shape Anthropic uses with Mythos — one underlying model, a safeguarded general release, and a restricted-access tier for vetted cybersecurity and life-sciences organizations. Two labs, arriving independently at the same answer within days of each other.

The industry has quietly converged on a conclusion: frontier capability will not be gated by refusing to build it, or by refusing to ship it. It will be gated by who gets the version with fewer brakes.

The pause, in retrospect

Twenty-five days separate the suspension from the launch. That is the number worth sitting with.

The optimistic reading is that the hold worked exactly as designed. Internal testing surfaced a capability that exceeded the deployment posture, the deployment stopped, safeguards and a vetted-access program were built, and the model shipped with a classification attached and a disclosure in progress. That is a governance process completing a cycle, in public, in under a month.

The skeptical reading is that a threshold you can clear in twenty-five days by adding refusals and an application form is a speed bump, not a gate. The capability did not change between August 9 and September 3. The packaging did.

Both readings are defensible from the outside, and neither can be settled without evaluation detail OpenAI has not published. What can be observed is that the Preparedness Framework's most severe designation did not, in its first real invocation, prevent shipping. It produced a tier structure.

What defenders actually get

For security teams the practical question is not the classification. It is throughput.

A model that saturates ExploitBench is, for the first time, plausibly capable of doing the expensive part of vulnerability research at machine speed — the part where a human spends days establishing whether a suspicious code path is reachable and weaponizable. Daybreak's stated use cases (validation, malware analysis, detection engineering) are exactly the tasks that are labor-bound rather than insight-bound.

The corresponding problem is that the same capability, with the same speed advantage, is now a known quantity to everyone who reads a launch post. Public Astra refuses PoC generation. Refusals are a filter, not a wall, and the last two years of jailbreak research is a fairly complete argument that determined adversaries route around them.

There is also the disclosure asymmetry to consider. Astra found two zero-days in testing and OpenAI is reporting them upstream. That is the right behavior. It is also a preview of a volume problem: if a model at this level is pointed at open-source infrastructure at scale, the finding rate will exceed the maintainer capacity to patch — a dynamic Flux has already seen in the open-weights model that surfaced 2,436 vulnerabilities and then waited two weeks for anyone to triage them.

The pricing footnote

Astra's reported launch pricing is $10 per million input tokens and $50 per million output, with cached reads at $1 per million, batch at half rate, and a Fast mode at 2x — roughly 2.5x the GPT-5.6 Sol rate. It carries a 1,050,000-token context window, 128K max output, and an April 30, 2026 knowledge cutoff.

That $10/$50 sticker is identical to the price Anthropic is charging for Fable 5.1, released two days earlier. The frontier tier has found its number. What separates the two now is not the rate card. It is what each lab will let you do at it.

#openai#gpt-6-astra#preparedness-framework#exploitbench#ai-security

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.