OpenAI’s Ultrafast preview runs GPT-5.6 Sol at up to 14×
By AgentRiot Editorial
Cerebras-backed, vendor-claimed 750 output tokens per second, limited preview. Standard Sol stays the public default.

OpenAI did not replace GPT-5.6 Sol on 13 August 2026. It added a speed tier.
The post is Previewing Ultrafast mode. Ultrafast is a new API service tier that OpenAI says runs GPT-5.6 Sol “up to 14× faster than Standard processing,” “up to 750 output tokens per second,” on Cerebras hardware. Access is a limited preview for a select group of customers. Everyone else gets a notify form.
Treat 14× and 750 tok/s as OpenAI-reported. The post does not publish a measurement method, a prompt set, or an independent clock.
What actually launched
Ultrafast is a processing mode for the existing Sol model, not a new checkpoint. The same prompt is supposed to produce the same kind of work, faster. OpenAI’s own demo is a side-by-side 3D warehouse simulator: Ultrafast Sol and standard Sol from one text prompt.
The company frames this as a break from the usual trade. Until now, real-time speed meant a smaller or more specialized model. Ultrafast is the claim that frontier Sol can sit in the latency slot.
That is only interesting if the preview numbers hold outside OpenAI’s demo reel. We do not have that evidence yet.
Where OpenAI says it is using it
The post lists five customer-shaped scenarios: incident response, financial research and security, customer support and voice, commerce, and interactive research loops that used to be overnight batch jobs. These are illustrations, not named customer case studies. No company is quoted with a measured before/after.
Internally, OpenAI says a group of its own developers is using Ultrafast on incident response — logs, traces, conversations, next checks — and on research loops that used to wait until morning. Engineers stay responsible for judgment and deployment. That is the company’s own usage note, not a third-party eval.
Cerebras is the hardware story
Ultrafast is also a partnership announcement. OpenAI says Cerebras is now serving its “most intelligent model” for this tier. The 750 tok/s figure is repeated in that paragraph. There is no public price, no rate-limit table, no region list, and no statement that ChatGPT consumers get the same mode.
If you already have Sol in the API, nothing in the post says your default traffic moved. The preview is opt-in and capacity-gated.
How this sits next to our Sol coverage
We have already written the Sol/Luna split, the Luna price cut, and the ChatGPT Instant/unlimited Luna routing. Those pieces were about which model you get and what it costs. This one is about a vendor-claimed speed multiplier on one of those models, for a subset of API customers, on someone else’s silicon.
Do not collapse them. Ultrafast does not change Luna pricing. It does not make Sol unlimited. It does not ship a 14× ChatGPT toggle.
What to do with it
If you are already an API customer with a latency-bound Sol workflow, the only concrete action in the post is the waitlist. If you are shopping models on tokens per second, ask OpenAI for the measurement setup before you treat 14× as a planning number. If you need the public Sol we already covered, keep using the standard tier until Ultrafast is generally available and independently timed.
Sources
- OpenAI, “Previewing Ultrafast mode” (13 August 2026)
- Ultrafast interest form

