SPXNDXDJIBTCETHOILGLD10YGOOGAAPLNVDATSLAMSFTMETASOLXRPLINKLTCDOTBNBSPXNDXDJIBTCETHOILGLD10YGOOGAAPLNVDATSLAMSFTMETASOLXRPLINKLTCDOTBNB
Home AI

When OpenAI Hit Its Own Critical Cyber Threshold, the Framework Logged It Rather Than Stopping It

OpenAI said on September 1 that GPT-6 Astra is the first model to cross the “Critical” cybersecurity threshold in its Preparedness Framework, the highest severity tier…

Two-colour risograph poster in coral and black on newsprint showing the OpenAI knot logo beside a padlock shackle dissolving into halftone dots

OpenAI said on September 1 that GPT-6 Astra is the first model to cross the “Critical” cybersecurity threshold in its Preparedness Framework, the highest severity tier the company defines for itself. Two days later it began rolling the model out.

Most of the coverage has treated the rating as the story. The sequence is the story. A severity ceiling that a model can hit and still ship 48 hours later is not a gate on deployment, it is a disclosure schedule wearing safety vocabulary. That distinction deserves saying plainly, because the whole policy case for letting frontier labs govern themselves rests on the claim that these frameworks are capable of returning a no.

What Critical Means in OpenAI’s Own Definition

The company’s published bar for Critical cyber capability is a model that can identify and develop functional zero-day exploits across many hardened real-world systems without human intervention, or devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal.

By OpenAI’s account, Astra reached it. The company describes the model as able to find previously unknown security flaws and build working exploits across well-protected systems without a person directing each step, and security trade coverage has reported benchmark results including a perfect score on exploit generation and the discovery of zero-day flaws during evaluation. OpenAI held back parts of Astra’s development and release for several weeks while it strengthened protections against cyber misuse, then concluded the safeguards sufficiently minimized the risk of severe harm.

That delay is genuine friction and deserves credit. It is also the entire extent of the constraint.

The Framework Describes. It Does Not Block.

Read the Preparedness Framework as an engineering document and it performed. A threshold was defined in advance, a model crossed it, the crossing was measured, a system card was published, the release was delayed, mitigations were added, and distribution was restricted. Every step executed.

Read it as a control and it fails at the only moment that counts, because there is no branch in which the answer is no. The most severe rating OpenAI can assign its own model produced a staged rollout. When the grader and the graded share a balance sheet, and the model in question is the company’s flagship commercial release, a finding that the safeguards are sufficient is not an independent finding. It is a business decision with a methodology section.

This is not a hypothetical governance concern. We have watched the same structure get drafted into law: the Great American AI Act would preempt state rules and require frontier developers to self-report safety results, which is precisely this arrangement with a federal stamp on it. The competing instinct showed up in the executive order seeking 30-day government access to frontier models before public release, an attempt to put somebody else in the room before the ship decision rather than after it. Astra is the case study that makes the sequencing question concrete.

Daybreak Puts OpenAI in the Licensing Business

The offensive capability is not general release. OpenAI routes it through Daybreak, an application-based program the company splits into two tiers:

  • Blue tier for defenders, covering secure code review, malware analysis and patch validation.
  • Red tier, limited to authorized vulnerability research and exploit testing.

Initial access goes to a small alpha group described as individuals and organizations responsible for protecting critical digital infrastructure, including the US government. Ordinary reasoning and coding capability reaches ChatGPT Plus, Pro, Business and Enterprise users, along with API and AWS customers, through normal channels.

Look at that structurally rather than as a product announcement. A private company has built and now operates an allowlist that determines which organizations receive automated zero-day discovery and exploit generation. For every comparable capability in the physical economy, from strong cryptography to munitions to dual-use manufacturing equipment, that judgment sits inside a government export-control regime with published criteria, an appeals path, and legislative oversight. Here it sits in a vendor’s application queue.

None of that is an accusation that OpenAI is running Daybreak carelessly. The evidence suggests the opposite. It is an observation that the criteria are unpublished, the appeal is a support ticket, and the entity deciding who is trustworthy is the same entity that books the revenue when the answer is yes. Anthropic, Google and Meta will reach this threshold within months, and on current form each will answer it privately too.

The Security Market’s New Chokepoint

The commercial consequence lands on cybersecurity vendors. Offensive security, penetration testing and vulnerability research are labor businesses whose margins depend on scarce human expertise. A model of this capability compresses that expertise into an API call, but only for the firms holding the right tier.

That produces a two-tier market with nothing to do with product quality. A mid-size penetration testing firm outside Daybreak now competes against Red-tier holders carrying a structural cost advantage it cannot hire, buy or engineer its way past. Its options narrow to acquisition or exit, and the buyers will be the firms already inside. Consolidation in security services is now partly a function of who OpenAI approved.

The defensive logic also has a hole in it. Attackers operating outside any allowlist will get comparable capability from open-weight models on a lag measured in months, not years. The allowlist meaningfully constrains the defenders who follow rules and mostly inconveniences the adversaries who do not.

What Should Happen Next

Our position is that OpenAI made a defensible call on the model and an indefensible one on the process. Shipping Astra with staged access, weeks of added safeguards and a restricted tier for offensive use is reasonable, and a blanket ban on capable security models would hand the field to people who ignore bans. But deciding unilaterally, against unpublished criteria and with no external review, which organizations receive state-grade offensive cyber capability is not a call a private company should be making alone. OpenAI should not want to own it either, because the liability of having approved the wrong applicant is close to unbounded.

The fix is neither prohibition nor the status quo. Publish the Daybreak criteria. Put an external body in the approval loop for Red tier, where CISA is the obvious candidate. Make a Critical designation trigger a mandatory pre-deployment review by someone who does not book the revenue. None of that prevents the model from shipping. All of it removes the part that should worry a business audience most, which is that the most consequential security decision of the year was made by the party with the strongest interest in making it quickly.