Smart Market InsightAI & SaaS Reviews
AI Tools

GPT-6 Astra: Why OpenAI Rated Its Own AI 'Critical' Risk

By Smart Market Insight EditorialPublished September 4, 20267 min read

Smart Market Insight Editorial

Editorial Team

Last verified: September 4, 2026

This article may contain affiliate links. We only recommend tools we’ve personally tested. Read our full disclaimer.

OpenAI began rolling out GPT-6 Astra on September 3, 2026, and told the world something no AI lab has said about its own flagship model before: this one crosses the "Critical" cybersecurity threshold in the company's own risk framework. That's not a marketing claim — it's a formal classification that triggers mandatory extra safeguards under OpenAI's Preparedness Framework, the same policy the company used two months ago to justify pausing parts of Astra's rollout in the first place.

If you use ChatGPT, Codex, or the OpenAI API for anything security-adjacent, this is one of the more consequential releases of the year, for reasons that have little to do with the usual benchmark chart. Here's what launched, how OpenAI is containing the risky part, and how it lines up against what Anthropic and Google shipped this same week.

Quick Take

  • What it is: GPT-6 Astra, OpenAI's new flagship model for ChatGPT, Codex, and the API, began a staged rollout on September 3, 2026, after OpenAI classified it as the first model to reach "Critical" cybersecurity capability under its Preparedness Framework.
  • What "Critical" means: Per OpenAI's own definition, a model at this level can independently find and build working zero-day exploits across hardened real-world systems, or plan and run a full cyberattack from a high-level goal with minimal human guidance.
  • The evidence: OpenAI says Astra scored a perfect result on its internal exploit-development benchmark and, during testing, autonomously found and chained working zero-day exploits without being specifically asked to.
  • How it's restricted: General chat, coding, and reasoning are available to all ChatGPT and API users on paid plans. Astra's offensive cybersecurity capability is walled off behind a new vetted tier, Daybreak Blue, and OpenAI has added monitoring that can pause or stop a task mid-run if it looks like misuse.
  • Pricing: $10 per million input tokens and $50 per million output tokens through the API, with a 1.05-million-token context window. It's not available on ChatGPT's free tier, and enterprise admins must turn it on per workspace.

What "Critical" Actually Means Here

OpenAI's Preparedness Framework ranks model risk across categories including cybersecurity, biological, and chemical harm, with four tiers running Low to Critical. No OpenAI model had been rated Critical in any category before this. According to OpenAI's announcement, Astra crossed that line on cybersecurity specifically: the ability to discover unknown vulnerabilities and turn them into functioning exploits against well-defended, real-world systems, largely on its own.

This isn't OpenAI's first brush with the issue. We covered GPT-5.6-Cyber in August, an offshoot of GPT-5.6 Sol gated behind a vetted-partner program called Daybreak Red after it hit "High" — one notch below Critical — and started surfacing real zero-days in Chrome and mobile operating systems. At the time, OpenAI disclosed it had paused parts of Astra's rollout on August 8 because testing couldn't rule out a Critical rating. That pause is now resolved, and the answer was yes.

How OpenAI Is Trying to Contain It

Astra isn't a locked-away research model — it's what now powers everyday ChatGPT and Codex sessions for paying users. OpenAI's answer is to split the model's behavior rather than the model itself: standard writing, coding, and agentic tasks work like any other release, available through ChatGPT Plus, Pro, Business, and Enterprise plus the API. The part that crossed Critical — offensive cybersecurity workflows like exploit development — sits behind Daybreak Blue, a new vetted tier for security teams doing defensive work, expanding gradually from an initial small test group.

OpenAI also added active monitoring that can interrupt a task mid-run. In ChatGPT or Codex, a flagged action can prompt you to confirm before continuing; in the API, a flagged task simply stops. OpenAI has acknowledged this system will sometimes misfire, flagging legitimate research or unrelated long-running work as suspicious — a real cost for developers, not just a footnote.

Pricing and Where You Can Get It

Astra's API pricing is $10 per million input tokens and $50 per million output tokens — the same rate OpenAI set for GPT-5.6 Sol after July's price cuts, with an optional Fast mode running roughly 2.5x faster at double the price. It carries a 1.05-million-token context window (about 922,000 input, 128,000 output).

Rollout is staged: a limited set of organizations got access first on September 3, with broader ChatGPT and API access, plus AWS, following over subsequent days. It's not on ChatGPT's free tier, and Enterprise admins must explicitly enable it per workspace rather than getting it by default — unusual friction for a flagship launch, and a clear signal of how seriously OpenAI is treating the classification.

The AGI Claim Nobody Asked For

Astra's launch came with more than a safety rating. OpenAI president Greg Brockman told reporters "welcome to the AGI era" and called the model a "generational leap," while leaving the actual determination up to "the reader." Sam Altman has separately said he expects OpenAI to reach a system it would internally call AGI by the end of 2026, while acknowledging it isn't there "quite yet."

Worth reading skeptically: AGI claims from the company selling the model are exactly what this site discounts until there's independent, reproducible evidence, and OpenAI's own framing here is notably hedged. The Critical cybersecurity rating, by contrast, is a concrete classification backed by a published benchmark result — the more useful thing to actually pay attention to this week.

How This Lines Up Against Claude and Gemini Right Now

Astra didn't launch into a quiet week. Anthropic shipped Claude Fable 5.1 on September 1, priced roughly 25% below Fable 5 with a 75% cut to cached-context pricing — a move about cost and agent economics, not a capability threshold. Google shipped Gemini 3.8 Flash the same week at the same introductory pricing as its predecessor ($0.75/$3.75 per million tokens), plus its own gated cybersecurity variant for government and infrastructure partners.

The pattern is worth noticing: rivals are competing on price and efficiency this week, while OpenAI is publicly grappling with a capability its own framework says needs restricting. That's not automatically a mark against Astra — a lab admitting a threshold was crossed and building controls around it is the system arguably working as designed. But anyone evaluating Astra for business use should weigh the access friction (approval gates, mid-task interruptions, no free tier) that Claude and Gemini don't currently carry, and note that regulatory scrutiny is rising alongside capability — the kind of thing the EU AI Act's content and risk-disclosure rules were built to watch for.

Frequently Asked Questions

Is GPT-6 Astra available to everyone right now? General chat, coding, and reasoning are rolling out to ChatGPT and API paid tiers over the days following the September 3 launch. It's not on the free tier, and the offensive capability that triggered the Critical rating is restricted to the vetted Daybreak Blue program.

What does a "Critical" cybersecurity rating actually mean? Under OpenAI's Preparedness Framework, it means the model can, largely without step-by-step human help, find unknown vulnerabilities and build working exploits against hardened real-world systems, or plan and carry out a cyberattack from a high-level goal.

Does this mean Astra can be used to hack things? The capability exists, which is why OpenAI gated it. Offensive workflows require Daybreak Blue approval, and monitoring can pause or stop tasks in ChatGPT, Codex, and the API that look like misuse.

Is Astra actually AGI? No independent body has confirmed that, and OpenAI's own executives left the label open to interpretation rather than declaring it outright. Treat it as framing around a strong model release, not a settled fact.

How does Astra's pricing compare to Claude and Gemini? At $10/$50 per million tokens, Astra costs more than Gemini 3.8 Flash ($0.75/$3.75) and sits in the same range as Claude Fable 5.1. Flash and Fable 5.1 target high-volume use; Astra's pricing reflects a flagship, higher-capability tier.

Bottom Line

GPT-6 Astra matters less for its benchmark scores than for what it forced OpenAI to admit: its own model now meets the bar the company set for an AI capable of independently finding and exploiting real-world vulnerabilities. The safeguards — Daybreak Blue gating, mid-task monitoring, no free-tier access — read as a serious response, but they add real friction for legitimate users. Evaluating Astra for coding or agentic work? Budget for occasional interruptions. Evaluating the AGI claims attached to it? Wait for evidence beyond the press cycle.

Related Articles