On the night of August 13, 2026, @OpenAI posted a three-tweet preview that got 1.7M views and immediately split reactions: Ultrafast mode, a Cerebras-powered inference tier that runs GPT-5.6 Sol at up to 750 tokens per second — as much as 14x the model's normal speed. It's the formal branding of a Cerebras partnership OpenAI first mentioned in passing during GPT-5.6's July 9 GA rollout, now shipping as a named preview product — to a short list of API customers, with no pricing attached.
The announcement lands one day after Gemini 3.7 Flash's launch put speed back at the center of frontier-model marketing, and independent Hacker News commentary on that same Gemini thread flagged the Ultrafast preview as a direct response. Here's what OpenAI actually announced, who gets it, what's missing, and how the reaction split between businesses that want in and subscribers who feel left out.
TL;DR — Ultrafast mode for GPT-5.6 Sol
| Question | Answer |
|---|---|
| What is it? | A high-speed inference tier for GPT-5.6 Sol running on Cerebras wafer-scale silicon |
| How fast? | Up to 750 tokens/second — up to 14x GPT-5.6 Sol's normal speed |
| Announced | August 13, 2026, 10:31 PM — three posts from @OpenAI, 1.7M views |
| Who gets it now? | A select group of API customers only |
| ChatGPT / Codex access? | Not announced. API-first launch; no consumer rollout confirmed |
| Pricing | Not disclosed — no rate card published as of this post |
| Hardware partner | Cerebras — same wafer-scale platform behind Gemma 4's 1,851 TPS preview |
| How to request access | Businesses can request notification via OpenAI's signup — no self-serve waitlist date given |
What Ultrafast mode actually is
OpenAI's second post in the thread carries the technical claim:
"Powered by @Cerebras, Ultrafast generates up to 750 tokens per second, bringing our most intelligent model to products and workflows where every second counts. Ultrafast is designed for businesses where faster frontier intelligence creates a measurable advantage..."
Two numbers matter here, and they describe the same thing from different angles. 750 tokens per second is the absolute throughput figure — the raw generation speed Cerebras's wafer-scale chips deliver for GPT-5.6 Sol. Up to 14x is the relative claim: how much faster that is than GPT-5.6 Sol's standard inference path on OpenAI's normal serving infrastructure.
This is not a new, smaller, or distilled model. OpenAI is explicit that Ultrafast runs the full GPT-5.6 Sol — "our most intelligent model" — just on different silicon. That distinction matters for anyone assuming speed tiers usually mean a capability trade-off, the way GPT-5.6 Luna trades some capability for cost. Ultrafast is explicitly positioned as no such trade — same intelligence, different latency.
This isn't the first mention
Readers who followed explainx.ai's July 9 GPT-5.6 GA coverage will recognize the number: that post already flagged "Cerebras Sol up to 750 tps in July (select customers)" as a line item in OpenAI's GA thread. The August 13 announcement is best read as OpenAI formalizing that effort with a product name — Ultrafast mode — a proper preview signup flow, and a headline 14x framing, rather than an entirely new capability appearing from nothing.
Who gets it now — and who doesn't
OpenAI's framing is unambiguous about scope. The first post:
"Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows."
That's three qualifiers in one sentence: preview, API only, select group. The third post reinforces it:
"We're working with an initial group of customers to understand where this speed makes the biggest difference, and how those learnings can inform our products over time. If your business requires frontier intelligence at the highest speed, you can request to be notified..."
There is no mention of ChatGPT or Codex anywhere in the thread. That's a notable omission given OpenAI's own August 6-7 ChatGPT update made GPT-5.6 Sol the default chat model for Plus/Pro users and unified Instant/reasoning behind one model — the natural next step would be surfacing Ultrafast in that same surface. It hasn't happened, and OpenAI hasn't said it will.
The pricing gap OpenAI hasn't addressed
None of the three announcement posts mention price. That's a real gap, not an oversight to gloss over. Hacker News commenter modeless, discussing the Ultrafast launch on the Gemini 3.7 Flash thread, put it plainly:
"It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far... Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though..."
GPT-5.6 Sol's standard API pricing is $5 input / $30 output per million tokens. Whether Ultrafast carries a premium over that rate — and how large — is genuinely unknown. Cerebras's wafer-scale chips are specialized hardware with different provisioning economics than commodity GPU clusters, and HN commenter piyh flagged the plausible outcome directly: "Sol on Cerebras is going to be expensive AF." A reply from jaggederest pushed back usefully, questioning whether wafer-scale inference is actually more or less expensive per token than GPU inference at scale — noting that Cerebras's pitch has always been throughput-per-dollar efficiency, not just raw speed, so the cost direction isn't obvious either way.
explainx.ai's read: don't trust any specific per-token number for Ultrafast you see circulating online right now — OpenAI hasn't published one. Treat pricing as an open question until a rate card appears, and budget for a premium over Sol's standard rate given the specialized hardware, without assuming how large.
Why 14x speed matters — real use cases
OpenAI's own framing names the target buyer directly: "businesses where faster frontier intelligence creates a measurable advantage." Three categories fit that description:
- Real-time agentic workflows. Multi-step agent loops — plan, call a tool, verify, retry — compound latency at every hop. At 750 TPS, a chain that took several seconds per step starts approaching the interactivity threshold, the same argument Cerebras and Google DeepMind made for Gemma 4's multimodal preview: "if every model was doing 2,000 tokens per second, you would build different products."
- Latency-sensitive business logic. Trading systems, fraud scoring, and other decision pipelines where a model's answer has to land inside a hard time budget benefit directly from throughput that removes queueing as the bottleneck — not just makes existing calls feel snappier.
- Customer-facing product experiences. Chat and voice interfaces where perceived responsiveness drives retention — every added second of "thinking" visibly costs engagement, and Ultrafast targets exactly that gap between frontier intelligence and consumer-grade latency.
The reaction: paying subscribers feel skipped
Two threads of criticism dominated the replies to OpenAI's announcement, and both are worth taking at face value rather than dismissing as noise.
"Why not Codex?"
MarsIQ (@SingleLayerch) voiced the sharpest version of subscriber frustration:
"why would this not be provided for codex users? Especially the ones paying 200 a month. This selective bs gets on my nerves. I pay 200 month for the best and earliest access to models."
The complaint tracks a real pattern: OpenAI's highest-paying consumer tier ($200/month ChatGPT Pro, which also carries Codex usage) does not automatically get first access to new capability previews when those previews launch API-first. Ultrafast follows that same path — businesses with API contracts, not individual power users, are first in line.
"Why announce it at all?"
omni (@omni1896837) raised the harder question:
"Why announce if it's not available to basically anyone?"
This is a fair "vaporware-adjacent" critique. A preview with a signup form and no committed general-availability date reads, to a skeptical audience, like marketing ahead of substance — a pattern OpenAI has drawn criticism for before with staged capability announcements. The counter-argument is that OpenAI explicitly frames this as a learning phase — the third post says the goal is working with early customers "to understand where this speed makes the biggest difference" before broader rollout — which is a legitimate reason to preview narrow before scaling wide. Both readings are defensible; the honest takeaway is that "preview" here means genuinely early, not imminent GA.
The competitive angle: Cerebras vs. Gemini 3.7 Flash's speed story
The timing is not incidental. Gemini 3.7 Flash launched August 13 with speed as a headline positioning point, and Hacker News commentary on that exact thread tied the two launches together directly. Commenter jjcm wrote:
"Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast..."
That's an independent, same-day observation — not OpenAI's own framing — that Ultrafast mode functions as a direct answer to Gemini's speed positioning. Whatever throughput advantage Gemini 3.7 Flash claims on Google's own infrastructure, GPT-5.6 Sol at 750 TPS on Cerebras is a materially different number, running the full frontier model rather than a smaller "flash" tier.
This also extends a broader pattern explainx.ai has tracked since June: Cerebras is becoming the inference speed layer for multiple labs simultaneously, not an exclusive partner to any one of them. Kimi, GLM, GPT-OSS, and Qwen already run on Cerebras hardware for open-weight speed. Google DeepMind added Gemma 4 31B multimodal at 1,851 TPS in June. OpenAI joining with a full frontier closed model — not just an open-weight release — is the clearest signal yet that wafer-scale silicon has become a competitive requirement at the frontier, not a niche optimization.
What to watch next
| If you are… | What to do |
|---|---|
| An API customer wanting early access | Request notification via OpenAI's Ultrafast signup — no committed timeline, but "expanded access to more businesses as capacity grows" is the stated direction |
| A Codex/ChatGPT Pro subscriber | Don't expect Ultrafast in your surface yet — nothing in OpenAI's posts commits to consumer access, despite the $200/month tier already carrying Codex usage |
| Budgeting for production use | Wait for a published rate card before modeling costs — treat any number you see elsewhere as unverified until OpenAI ships pricing |
| Comparing against Gemini 3.7 Flash | Read both launches together — see the full Gemini 3.7 Flash comparison for where each model's speed claim actually comes from |
| Tracking the Cerebras partnership pattern | See Gemma 4 31B on Cerebras for how the same wafer-scale platform served Google DeepMind two months earlier |
Related reading
- GPT-5.6 Sol, Terra, and Luna: OpenAI Preview Launch Explained — original July 9 GA thread that first mentioned Cerebras at 750 TPS
- GPT-5.6 Sol Now Runs All of ChatGPT — Free Users Get Unlimited Chats — August 6-7 consumer-surface update, no Ultrafast mention
- Gemma 4 31B on Cerebras: 1,800+ TPS — Cerebras's prior frontier-lab partnership, June 2026
- Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 comparison — the same-day launch Ultrafast is competing against
- GPT-5.6 vs Claude Fable 5 comparison — full benchmark matrix for GPT-5.6 Sol
- Castform + Neon: open model matches Sol on retrieval at ~100x lower cost — cost-efficiency counterpoint from the open-weight side
- OpenAI cuts Luna 80%, Terra 20% pricing — GPT-5.6 family pricing history ahead of Ultrafast
Official sources: @OpenAI on X, August 13, 2026 · Cerebras blog: Accelerating GPT-5.6 Sol Ultrafast · Hacker News discussion on the Gemini 3.7 Flash thread (662 points), commenters jjcm, modeless, piyh, and jaggederest
Speed, access, and pricing details reflect OpenAI's August 13, 2026 preview announcement. Ultrafast mode is an early preview with no committed general-availability date or published pricing — verify current access and rates on openai.com before making production decisions.
