GPT-6 Astra is working in Pro and the API. Plus is still waiting. Fable 5.1 still leads the independent index.
By AgentRiot Editorial
Two days after launch, OpenAI has GPT-6 Astra live for Pro, Enterprise, Business Premium, and the API. Plus is not done. Independent Artificial Analysis v4.2 still ranks Claude Fable 5.1 first. The vendor table is now complete, and playable Astra games are up.

GPT-6 Astra is not a rumor and it is not a gated preview for the people OpenAI named on September 4. @OpenAI wrote that it is available to all Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex, and live in the API. A follow-up the same evening said that in Chat, Astra powers GPT-6 Pro for Pro, Business, and Enterprise.
It is also not fully rolled out. That same post said Plus and Business “might take a few days.” As of September 5, that is still the official line. If your model picker is empty, that is the queue, not a dead model.
This is a two-day follow-up to AgentRiot’s launch write-up. The September 3 piece had a mixed vendor scorecard and no independent Artificial Analysis row. Both of those facts moved.
What “live” actually means today
The public name is still GPT-6 Astra. The API ID is still gpt-6-astra. The model card still lists a 1,050,000-token context window, 922,000 max input, 128,000 max output, an April 30, 2026 knowledge cutoff, and reasoning.effort of low, medium, high, xhigh, and max.
Standard API rates on that card and the pricing page are unchanged: $10 per million input, $1 cached input, $12.50 cache writes, $50 output. Prompts over 272K tokens are $20 / $2 / $25 / $75 for the full request. Fast mode is 2× those rates. Batch and Flex are 50%. A synthetic 1M input + 1M output job with no cache and no tools is $60.
That is the same uncached list as Claude Fable 5.1 in AgentRiot’s launch comparison, and 2.5× GPT-5.6 Sol on that workload ($4 + $20 = $24). Fast mode is unavailable for Astra with EU data residency.
Where it is on: Pro / Enterprise / Business Premium in Work and Codex, GPT-6 Pro in Chat for Pro / Business / Enterprise, the OpenAI API, and the original launch promise of Azure and AWS Bedrock. Enterprise still had to be turned on by an admin at launch; that default-off note has not been publicly reversed.
Where it is not: a completed Plus rollout. Computer-use still wants the ChatGPT desktop app. Default Astra still refuses proof-of-concept exploit work. Daybreak is the later, looser defensive path. “Fully functional” is true for the surfaces that have the model. It is not true as a synonym for “everyone has it” or “every cyber workflow is open.”
The independent scoreboard finally exists, and Fable 5.1 still leads it
AgentRiot’s launch article said Artificial Analysis had no GPT-6 Astra page. That page exists now.
On September 4, @ArtificialAnlys shipped Intelligence Index v4.2: AA-Briefcase and Surge AI’s GDP.pdf in, GPQA Diamond out, held-out tests raised to 40% of the index. Absolute v4.2 numbers are not comparable to the v4.1.1 row OpenAI printed at launch (Astra 61.2, Fable 5.1 65.7).
What Artificial Analysis stated in text, without relying on chart pixels:
- Claude Fable 5.1 leads the Index. GPT-6 Astra is second, with a 4-point gain over GPT-5.6 Sol.
- Astra “dominates the output token Pareto frontier” near the intelligence frontier. Fable 5.1 and Gemini 3.8 Flash sit at the high-token end among models scoring at least 25.
- On AA-Briefcase, Fable 5.1 and Opus 5 lead, then Astra and Muse Spark 1.3. Astra is about 85 Elo above Sol.
- On GDP.pdf all-pass, Astra 33.2%, Sol 28.2%, Fable 5.1 26.2%.
A public LMSYS Arena ranking for Astra was not available on September 5.
The independent conclusion matches the vendor one in shape: Astra is a real step over Sol, especially on documents and token use. It is not the Index leader.
Every row OpenAI printed
All scores below are OpenAI-reported from the launch post. Effort is “maximum at any effort.” GPT runs were in OpenAI’s research environment or API, which the company says can differ from production ChatGPT. Hyphens are missing cells in that table, not zeros. Do not treat this as an independent reprint.
AgentRiot’s launch piece printed 57.7% for Terminal-Bench 4.0. The live OpenAI table is 57.9%. Humanity’s Last Exam in that table is Astra 57.2%, Sol not printed, Fable 5.1 65.0%, Fable 5 63.8%, Opus 5 63.6%. Fable 5.1 still wins that printed row.
Computer use
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| Agents' Last Exam | 59.3% | 53.6% | - | 48.7% | 55.5% | - |
| OSWorld 2.0 (offline set, partial) | 72.6% | 65.7% | - | - | 70.2% | - |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | - | 87.3% | - | - |
OSWorld is also the efficiency claim versus Sol: 72.6% in about 40 minutes against 65.7% in about 75 minutes. Mind2Web, with the updated Codex harness, is the 1.9× faster-than-Sol computer-use number.
Professional
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | - |
| BenchCAD | 95.9% | 83.3% | 84.3% | 67.5% | 82.1% | - |
| BrowseComp | 91.5% | 90.4% | - | 87.4% | 90.8% | - |
| OpenScore String Quartets (1 - OMR-NED) | 0.84 | 0.19 | - | - | - | - |
| Internal Design Tasks | 50.0% | 47.4% | - | 35.8% | - | - |
| Internal Data Science Tasks | 40.9% | 30.5% | - | 34.7% | - | - |
| AA Intelligence Index v4.1.1 (vendor-printed) | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
Coding
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 44.5% | 52.6% | 19.1% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% |
| FrontierCode 1.1 Main | 53.3% | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% |
| Internal Database Migration Tasks | 63.9% | 42.7% | 57.8% | 50.3% | - | - |
| AA Coding Agent Index v1.4 (vendor-printed) | 67.0 | 65.1 | - | 67.2 | 68.1 | 61.2 |
Astra takes Terminal-Bench 4.0 and DeepSWE. It does not take FrontierCode Main or Extended. Fable 5 does, by a hair. The vendor-printed Coding Agent Index has Opus 5 at 68.1 against Astra 67.0.
Academic, science, health
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | - |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | - |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% |
| Humanity's Last Exam (w/ tools) | 57.2% | - | 65.0% | 63.8% | 63.6% | - |
| GeneBench Pro | 37.1% | 32.3% | - | - | - | - |
| MedChemBench (internal) | 49.3% | 47.4% | - | - | - | - |
| LifeSciBench | 60.3% | 59.9% | - | - | - | - |
| HealthBench Professional (length-adjusted) | 63.4% | 60.5% | 58.1% | 60.9% | 56.4% | 52.1% |
The blog lede still says FrontierMath “98%.” Use 97.6% from the table. OpenAI also published two prime-gap proofs with Astra in the footnotes, which is a research claim, not a leaderboard row.
Cyber, alignment, long context, ARC
Cyber numbers are vendor-reported, often without production safeguards on the evals. They are not instructions.
| Bench | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| ExploitBench | 100.0% | 78.5% | - | - | 70% | - |
| ExploitGym | 42.4% | 30.3% | 30.4% | 28.4% | 22.0% | - |
| ExploitBench (June–Aug 2026) | 39.0% | 5.5% | - | - | - | - |
| SRE-Bench (one attempt) | 88.0% | 55.9% | - | - | 12.5% | - |
| SEC-Bench Pro | 85.4% | 79.1% | - | - | - | - |
| Internal computer-use safety (lower better) | 2.4% | 22.0% | 9.5% | 18.3% | 11.5% | - |
| Same, with AutoReview (lower better) | 1.8% | 4.3% | - | - | - | - |
| Internal circumvention (lower better) | 0.00% | 0.29% | - | - | - | - |
| ExploitGym honeypot (lower better) | 0.0% | 48.2% | - | - | - | - |
| Impossible ExploitGym | 100.0% | - | - | - | - | - |
| Internal hallucination (lower better) | 4.2% | 12.2% | - | - | - | - |
| MRCR v2 8-needle 256K–512K | 100.0% | 91.5% | - | - | - | - |
| MRCR v2 8-needle 512K–1M | 96.3% | 73.8% | - | - | - | - |
| ARC-AGI-3 | 99.9% | 7.8% | - | - | 30.2% | - |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 89.2% | 90.4% | - |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 98.5% | 97.5% | - |
SRE-Bench four-attempt: Astra 99.2% vs Sol 68.7%. ARC-AGI-3 used OpenAI’s responses API harness. FrontierCode used a Codex-like developer message. ExploitGym dropped the 6-hour limit. Those footnotes matter if you are going to quote the highs.
The losses that survive the full sheet: Fable 5.1 on the vendor-printed Intelligence Index and on Humanity’s Last Exam as printed; Fable 5 on both FrontierCode headlines; Opus 5 on the vendor-printed Coding Agent Index. Astra’s blowouts are computer use versus Sol, FrontierMath T4, ARC-AGI-3, Terminal-Bench Science, AutomationBench, BenchCAD, and the cyber suite it chose to run without production safeguards.
The games are the part you can click
OpenAI’s product story was never the Index. It was “anything you can do on a computer.” Sites is in public beta on Plus, Pro, Business, Enterprise, and Edu, with plan limits. It creates, hosts, and shares websites, apps, and games from a prompt.
The OpenAI showcase now has Astra-tagged playable builds. These are the ones labeled GPT-6 Astra with a live link:
- Velocity Loop, a 3D toy-car time trial in miniature workshops, by VB Srivastav
- Hollowflux, a procedural dungeon crawler with reactive water, by Thomas Ricouard
- Little Ritual, coffee delivery on a small spherical world, by Jeff Wang
- Sunwake, a sailing game, by Thomas Ricouard
- Void Explorer, procedural planets you can land on, by Thomas Ricouard
Those are OpenAI staff demos, not a random X montage. Treat them as existence proofs that Astra-plus-Sites can ship a browser game, not as a quality ranking against a human studio.
Outside the showcase, the launch post itself credits Pietro Schirano. His September 3 thread is the clip reel: an underwater exploration game from one /goal prompt, a PS1-style Beyblade game, a jet-ski game with underwater reef, plus Blender-via-code and an Ableton MCP track. Those are videos on X, not hosted cartridges.
A hosted one that is: P(DOOM), a Doom 64-style AI-safety parody. Travis Fischer says Astra built it. The page itself does not mention GPT-6. Attribute the claim to him.
@OpenAIDevs is running a 24-hour “drop a demo” thread as of September 5. That stream will fill with more links, and with more junk. Prefer a live URL plus a named author over a screenshot of a menu.
Conclusion
Astra is a working frontier model on the surfaces OpenAI turned on September 4. People are already shipping playable 3D games from it. The API ID, the $10/$50 tariff, and the long context window are real.
It is not “fully rolled out.” Plus is still in the “few days” bucket. Default Astra still will not do exploit PoCs. Independent Artificial Analysis put it second to Fable 5.1 on Index v4.2, while giving it the document all-pass and the token-efficiency story.
If you are buying computer use, Sites games, long math, or fewer tokens per hard task, this is the model to try this weekend, on Pro or the API. If you are buying a single Intelligence Index number, cheap cache reads, or Plus-plan access today, wait or stay on Fable 5.1. OpenAI’s slogan was “most intelligent.” Two days of evidence, including OpenAI’s own full table, is more specific than that.

