GPT-6 Astra is live. OpenAI calls it the most intelligent model. The vendor table is more complicated.
By AgentRiot Editorial
OpenAI launched GPT-6 Astra on September 3 with SOTA computer-use and math scores, $10/$50 API pricing, and a mixed vendor scorecard against Claude Fable 5.1.

OpenAI launched GPT-6 Astra on September 3, 2026. The company calls it “the world’s most intelligent and aligned model.” That sentence is OpenAI’s. The same post publishes a vendor scorecard that does not uniformly beat Claude Fable 5.1.
The public name is GPT-6 Astra. The API ID is gpt-6-astra. @OpenAI led with a computer-use demo: “Anything you can do on a computer, Astra can do for you. Fast.” @AndrewGinns pointed at the blog a few minutes earlier. The blog is the source of record for what shipped.
This is the launch. It is not a rewrite of AgentRiot’s earlier Critical-cyber designation piece. That post was a policy preview. This one has a model ID, a price, a rollout, and a benchmark table.
What actually shipped today
Astra is rolling out first to a limited set of organizations. Over the coming days OpenAI says it reaches all ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API and AWS. Subscription usage sits inside existing allowances, with extra credits for sale. Pro, Business, and Enterprise also get GPT-6 Astra Pro. Enterprise admins must turn it on; default is off.
The model page lists:
- 1,050,000-token context window
- 922,000 max input tokens
- 128,000 max output tokens
- April 30, 2026 knowledge cutoff
- text and image in, text out
reasoning.effort:low,medium,high,xhigh,max
Responses and Chat Completions are supported. Batch is supported. Realtime, Assistants, fine-tuning, embeddings, and native image/video/speech endpoints are not. Tools on Responses include computer use, hosted shell, apply_patch, MCP, web search, code interpreter, and skills.
OpenAI wants you in the ChatGPT desktop app for the computer-use pitch.
Price: same list as Fable 5.1, more than Sol, worse cache than Fable
Standard API rates from the model card and pricing page, short context:
| Model | Input / 1M | Cached input / 1M | Cache writes / 1M | Output / 1M |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $12.50 | $50.00 |
| Claude Fable 5.1 | $10.00 | $0.25 | $12.50 (5-min) | $50.00 |
| GPT-5.6 Sol | $4.00 | $0.40 | $5.00 | $20.00 |
A synthetic 1M input + 1M output job, no cache, no tools: Astra $60, Fable 5.1 $60, Sol $24. Astra is 2.5× Sol on that workload. Prompts over 272K tokens jump to $20 / $2 / $25 / $75 for the full request. Fast mode is 2× those rates for up to 2.5× speed. Batch and Flex are 50% of Standard.
The rumor that Astra is “cheaper than Fable 5.1” does not survive the tariff. List input and output match. Cache reads do not: Astra $1 versus Fable 5.1’s $0.25. Token efficiency is a different claim, and OpenAI’s own wording is against GPT-5.6 Sol, not Anthropic.
The vendor benches, including the losses
All scores below are OpenAI-reported from the launch post. Effort is “maximum at any effort.” GPT runs were in OpenAI’s research environment or API, which the company says can differ from production ChatGPT. Independent Artificial Analysis does not have a GPT-6 Astra model page yet (404 on check). Treat this as a vendor scorecard, not a third-party leaderboard reprint.
Coding (six-column rows in the post: Astra, Sol, Fable 5.1, Fable 5, Opus 5, Gemini 3.8 Flash):
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% | 52.3% | 19.1% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% |
| FrontierCode 1.1 Main | 53.3% | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% |
Astra takes Terminal-Bench 4.0 and DeepSWE. It does not take either FrontierCode headline row. Fable 5 does, by a hair.
Academic / reasoning:
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% |
| Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
OpenAI’s own table puts Fable 5.1 first on the Artificial Analysis Intelligence Index it chose to print. Humanity’s Last Exam (with tools) is listed as Astra 57.2%, Sol 65.0%, Fable 5.1 63.8%, Fable 5 63.6%. That row is incomplete in the extract, but the four printed numbers already put Astra last among those four. If you came here because X said Astra is simply “better than Fable 5.1,” that is not what this table says.
Where Astra does blow the doors off, the gaps are large and still vendor-reported:
- FrontierMath Tier 4 (v2): 97.6% vs Sol 83.0% and Fable 5.1 87.8%. The blog also says Astra “saturates FrontierMath Tier 4 with a 98% score” in the lede. Use 97.6% from the table.
- ARC-AGI-3: 99.9% vs Sol 7.8%. OpenAI calls that saturation.
- Terminal-Bench Science 0.1: 64.6% vs Sol 22.4% and Fable 5.1 52.6%.
- AutomationBench: 41.4% vs Sol 18.1% and Fable 5.1 31.4%.
- ScreenSpot-Pro, no tools: 92.7% vs Sol 76.9%.
- OSWorld 2.0 (offline set, partial score): 72.6% in about 40 minutes vs Sol 65.7% in about 75 minutes — higher score, about 47% less time. That is the official token/time-efficiency claim against Sol.
@OpenAI also named Agents’ Last Exam, HealthBench Pro, and Terminal-Bench 4.0 as SOTA. The Agents’ Last Exam row in the post is truncated, so only Astra’s 59.3% is used here.
Computer use is the product, not a sidebar
The launch is built around doing work on a computer: forms, CRM updates, calendars, research into email or docs, plots, websites, frontend QA, installs, on-screen troubleshooting. Codex’s harness is updated alongside the model. OpenAI says that plus Astra’s efficiency is 1.9× faster than current GPT-5.6 Sol on Mind2Web.
In Codex, Astra can keep notes across context windows instead of compacting everything into one summary, and earlier windows stay searchable. OpenAI calls that experimental in config.toml, default “in the coming weeks.”
Sites in ChatGPT can create, host, and share websites, apps, and games from a prompt. Slide/doc work is pitched as template-faithful rather than “a deck that looks like a deck.”
None of that is independently timed outside OpenAI’s post. It is the product story they want you to buy.
Cyber: Critical, gated, and not a recipe book
Astra still meets the Critical cybersecurity threshold described in Path to Astra. Default Astra, launching today, is allowed to do secure code review and patching. OpenAI says it will refuse more advanced tasks such as creating proof-of-concept exploits. Daybreak is the path to looser defensive workflows later.
Vendor cyber numbers, OpenAI-reported, without production safeguards on the evals:
- ExploitBench: Astra 100%, Sol 78.5%
- ExploitGym: Astra 42.4%, Sol 30.3%, “substantially fewer output tokens”
- Internal ExploitBench (June–August 2026): Astra 39.0%, Sol 5.5%
- SRE-Bench, one attempt: Astra 88.0%, Sol 55.9% (four-attempt: 99.2% vs 68.7%)
OpenAI also says Astra found two zero-days during that internal port and is disclosing them. Expert-led tests without production safeguards found working chains against hardened browsers and operating systems. That is the company’s eval narrative. It is not a how-to.
The system card is up. One finding OpenAI put in the launch post: written reasoning is harder to monitor than GPT-5.6 Sol’s, because Astra writes fewer steps on simpler tasks. Alignment evals go the other way: an ExploitGym-style honeypot that Sol failed 48% of the time without production safeguards, Astra failed 0% of the time in the launch lede (the table prints 48.2% for Sol). Internal circumvention: Astra 0.00% vs Sol 0.29%. Hallucination bench (lower better): 4.2% vs 12.2%.
Misalignment monitoring can pause ChatGPT/Codex tasks for a human click, and stop API tasks outright. OpenAI warns that will catch some legitimate long-running work.
What “most powerful ever” can and cannot mean today
OpenAI’s claim is superlative. The evidence it published is domain-shaped.
Astra is the vendor leader on computer-use speed/accuracy against Sol, on FrontierMath Tier 4, ARC-AGI-3, Terminal-Bench Science, AutomationBench, GPQA Diamond, and the cyber suite it chose to print. It is not the vendor leader on the Artificial Analysis Intelligence Index in that same post, not on Humanity’s Last Exam as printed, and not on FrontierCode 1.1 Main or Extended.
Cost: same uncached list as Fable 5.1, four times Fable’s cache-read rate, 2.5× Sol on a raw 1M+1M job. Efficiency: OpenAI’s OSWorld clock and ExploitGym token note are vs Sol, in OpenAI’s harness, often with extra access flags on the cyber side.
Independent public leaderboards checked at publication time do not list GPT-6 Astra yet. Until Artificial Analysis and Agent Arena reprint it, the “most powerful model ever” line is a company slogan sitting on a mixed, self-run table.
If you are buying computer use and long math, this is the launch to watch this week. If you are buying a single Intelligence Index number or cheap cached agent loops, Fable 5.1 is still on the page OpenAI just posted.

