explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — Ultrafast mode for GPT-5.6 Sol
  • What Ultrafast mode actually is
  • Who gets it now — and who doesn't
  • The pricing gap OpenAI hasn't addressed
  • Why 14x speed matters — real use cases
  • The reaction: paying subscribers feel skipped
  • The competitive angle: Cerebras vs. Gemini 3.7 Flash's speed story
  • What to watch next
  • Related reading
← Back to blog

explainx / blog

GPT-5.6 Sol Ultrafast Mode: 750 Tokens/Sec via Cerebras, No Pricing Yet

OpenAI previewed Ultrafast mode for GPT-5.6 Sol on August 13, 2026 — up to 750 tokens/sec via Cerebras, 14x normal speed. Select API customers only, no pricing disclosed, no Codex/ChatGPT access announced.

Aug 14, 2026·10 min read·Yash Thakker
OpenAIGPT-5.6CerebrasInferenceAI Benchmarks
go deep
GPT-5.6 Sol Ultrafast Mode: 750 Tokens/Sec via Cerebras, No Pricing Yet

On the night of August 13, 2026, @OpenAI posted a three-tweet preview that got 1.7M views and immediately split reactions: Ultrafast mode, a Cerebras-powered inference tier that runs GPT-5.6 Sol at up to 750 tokens per second — as much as 14x the model's normal speed. It's the formal branding of a Cerebras partnership OpenAI first mentioned in passing during GPT-5.6's July 9 GA rollout, now shipping as a named preview product — to a short list of API customers, with no pricing attached.

The announcement lands one day after Gemini 3.7 Flash's launch put speed back at the center of frontier-model marketing, and independent Hacker News commentary on that same Gemini thread flagged the Ultrafast preview as a direct response. Here's what OpenAI actually announced, who gets it, what's missing, and how the reaction split between businesses that want in and subscribers who feel left out.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR — Ultrafast mode for GPT-5.6 Sol

QuestionAnswer
What is it?A high-speed inference tier for GPT-5.6 Sol running on Cerebras wafer-scale silicon
How fast?Up to 750 tokens/second — up to 14x GPT-5.6 Sol's normal speed
AnnouncedAugust 13, 2026, 10:31 PM — three posts from @OpenAI, 1.7M views
Who gets it now?A select group of API customers only
ChatGPT / Codex access?Not announced. API-first launch; no consumer rollout confirmed
PricingNot disclosed — no rate card published as of this post
Hardware partnerCerebras — same wafer-scale platform behind Gemma 4's 1,851 TPS preview
How to request accessBusinesses can request notification via OpenAI's signup — no self-serve waitlist date given

What Ultrafast mode actually is

OpenAI's second post in the thread carries the technical claim:

"Powered by @Cerebras, Ultrafast generates up to 750 tokens per second, bringing our most intelligent model to products and workflows where every second counts. Ultrafast is designed for businesses where faster frontier intelligence creates a measurable advantage..."

Two numbers matter here, and they describe the same thing from different angles. 750 tokens per second is the absolute throughput figure — the raw generation speed Cerebras's wafer-scale chips deliver for GPT-5.6 Sol. Up to 14x is the relative claim: how much faster that is than GPT-5.6 Sol's standard inference path on OpenAI's normal serving infrastructure.

This is not a new, smaller, or distilled model. OpenAI is explicit that Ultrafast runs the full GPT-5.6 Sol — "our most intelligent model" — just on different silicon. That distinction matters for anyone assuming speed tiers usually mean a capability trade-off, the way GPT-5.6 Luna trades some capability for cost. Ultrafast is explicitly positioned as no such trade — same intelligence, different latency.

This isn't the first mention

Readers who followed explainx.ai's July 9 GPT-5.6 GA coverage will recognize the number: that post already flagged "Cerebras Sol up to 750 tps in July (select customers)" as a line item in OpenAI's GA thread. The August 13 announcement is best read as OpenAI formalizing that effort with a product name — Ultrafast mode — a proper preview signup flow, and a headline 14x framing, rather than an entirely new capability appearing from nothing.


Who gets it now — and who doesn't

OpenAI's framing is unambiguous about scope. The first post:

"Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows."

That's three qualifiers in one sentence: preview, API only, select group. The third post reinforces it:

"We're working with an initial group of customers to understand where this speed makes the biggest difference, and how those learnings can inform our products over time. If your business requires frontier intelligence at the highest speed, you can request to be notified..."

There is no mention of ChatGPT or Codex anywhere in the thread. That's a notable omission given OpenAI's own August 6-7 ChatGPT update made GPT-5.6 Sol the default chat model for Plus/Pro users and unified Instant/reasoning behind one model — the natural next step would be surfacing Ultrafast in that same surface. It hasn't happened, and OpenAI hasn't said it will.


The pricing gap OpenAI hasn't addressed

None of the three announcement posts mention price. That's a real gap, not an oversight to gloss over. Hacker News commenter modeless, discussing the Ultrafast launch on the Gemini 3.7 Flash thread, put it plainly:

"It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far... Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though..."

GPT-5.6 Sol's standard API pricing is $5 input / $30 output per million tokens. Whether Ultrafast carries a premium over that rate — and how large — is genuinely unknown. Cerebras's wafer-scale chips are specialized hardware with different provisioning economics than commodity GPU clusters, and HN commenter piyh flagged the plausible outcome directly: "Sol on Cerebras is going to be expensive AF." A reply from jaggederest pushed back usefully, questioning whether wafer-scale inference is actually more or less expensive per token than GPU inference at scale — noting that Cerebras's pitch has always been throughput-per-dollar efficiency, not just raw speed, so the cost direction isn't obvious either way.

explainx.ai's read: don't trust any specific per-token number for Ultrafast you see circulating online right now — OpenAI hasn't published one. Treat pricing as an open question until a rate card appears, and budget for a premium over Sol's standard rate given the specialized hardware, without assuming how large.


Why 14x speed matters — real use cases

OpenAI's own framing names the target buyer directly: "businesses where faster frontier intelligence creates a measurable advantage." Three categories fit that description:

  • Real-time agentic workflows. Multi-step agent loops — plan, call a tool, verify, retry — compound latency at every hop. At 750 TPS, a chain that took several seconds per step starts approaching the interactivity threshold, the same argument Cerebras and Google DeepMind made for Gemma 4's multimodal preview: "if every model was doing 2,000 tokens per second, you would build different products."
  • Latency-sensitive business logic. Trading systems, fraud scoring, and other decision pipelines where a model's answer has to land inside a hard time budget benefit directly from throughput that removes queueing as the bottleneck — not just makes existing calls feel snappier.
  • Customer-facing product experiences. Chat and voice interfaces where perceived responsiveness drives retention — every added second of "thinking" visibly costs engagement, and Ultrafast targets exactly that gap between frontier intelligence and consumer-grade latency.

The reaction: paying subscribers feel skipped

Two threads of criticism dominated the replies to OpenAI's announcement, and both are worth taking at face value rather than dismissing as noise.

"Why not Codex?"

MarsIQ (@SingleLayerch) voiced the sharpest version of subscriber frustration:

"why would this not be provided for codex users? Especially the ones paying 200 a month. This selective bs gets on my nerves. I pay 200 month for the best and earliest access to models."

The complaint tracks a real pattern: OpenAI's highest-paying consumer tier ($200/month ChatGPT Pro, which also carries Codex usage) does not automatically get first access to new capability previews when those previews launch API-first. Ultrafast follows that same path — businesses with API contracts, not individual power users, are first in line.

"Why announce it at all?"

omni (@omni1896837) raised the harder question:

"Why announce if it's not available to basically anyone?"

This is a fair "vaporware-adjacent" critique. A preview with a signup form and no committed general-availability date reads, to a skeptical audience, like marketing ahead of substance — a pattern OpenAI has drawn criticism for before with staged capability announcements. The counter-argument is that OpenAI explicitly frames this as a learning phase — the third post says the goal is working with early customers "to understand where this speed makes the biggest difference" before broader rollout — which is a legitimate reason to preview narrow before scaling wide. Both readings are defensible; the honest takeaway is that "preview" here means genuinely early, not imminent GA.


The competitive angle: Cerebras vs. Gemini 3.7 Flash's speed story

The timing is not incidental. Gemini 3.7 Flash launched August 13 with speed as a headline positioning point, and Hacker News commentary on that exact thread tied the two launches together directly. Commenter jjcm wrote:

"Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast..."

That's an independent, same-day observation — not OpenAI's own framing — that Ultrafast mode functions as a direct answer to Gemini's speed positioning. Whatever throughput advantage Gemini 3.7 Flash claims on Google's own infrastructure, GPT-5.6 Sol at 750 TPS on Cerebras is a materially different number, running the full frontier model rather than a smaller "flash" tier.

This also extends a broader pattern explainx.ai has tracked since June: Cerebras is becoming the inference speed layer for multiple labs simultaneously, not an exclusive partner to any one of them. Kimi, GLM, GPT-OSS, and Qwen already run on Cerebras hardware for open-weight speed. Google DeepMind added Gemma 4 31B multimodal at 1,851 TPS in June. OpenAI joining with a full frontier closed model — not just an open-weight release — is the clearest signal yet that wafer-scale silicon has become a competitive requirement at the frontier, not a niche optimization.


What to watch next

If you are…What to do
An API customer wanting early accessRequest notification via OpenAI's Ultrafast signup — no committed timeline, but "expanded access to more businesses as capacity grows" is the stated direction
A Codex/ChatGPT Pro subscriberDon't expect Ultrafast in your surface yet — nothing in OpenAI's posts commits to consumer access, despite the $200/month tier already carrying Codex usage
Budgeting for production useWait for a published rate card before modeling costs — treat any number you see elsewhere as unverified until OpenAI ships pricing
Comparing against Gemini 3.7 FlashRead both launches together — see the full Gemini 3.7 Flash comparison for where each model's speed claim actually comes from
Tracking the Cerebras partnership patternSee Gemma 4 31B on Cerebras for how the same wafer-scale platform served Google DeepMind two months earlier

Related reading

  • GPT-5.6 Sol, Terra, and Luna: OpenAI Preview Launch Explained — original July 9 GA thread that first mentioned Cerebras at 750 TPS
  • GPT-5.6 Sol Now Runs All of ChatGPT — Free Users Get Unlimited Chats — August 6-7 consumer-surface update, no Ultrafast mention
  • Gemma 4 31B on Cerebras: 1,800+ TPS — Cerebras's prior frontier-lab partnership, June 2026
  • Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 comparison — the same-day launch Ultrafast is competing against
  • GPT-5.6 vs Claude Fable 5 comparison — full benchmark matrix for GPT-5.6 Sol
  • Castform + Neon: open model matches Sol on retrieval at ~100x lower cost — cost-efficiency counterpoint from the open-weight side
  • OpenAI cuts Luna 80%, Terra 20% pricing — GPT-5.6 family pricing history ahead of Ultrafast

Official sources: @OpenAI on X, August 13, 2026 · Cerebras blog: Accelerating GPT-5.6 Sol Ultrafast · Hacker News discussion on the Gemini 3.7 Flash thread (662 points), commenters jjcm, modeless, piyh, and jaggederest


Speed, access, and pricing details reflect OpenAI's August 13, 2026 preview announcement. Ultrafast mode is an early preview with no committed general-availability date or published pricing — verify current access and rates on openai.com before making production decisions.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 7, 2026

GPT-5.6 Sol Now Runs All of ChatGPT — Free Users Get Unlimited Chats

On August 6, 2026, OpenAI folded ChatGPT's separate Instant and reasoning models into one GPT-5.6 Sol experience for Plus and Pro, and rolled out unlimited text chats on GPT-5.6 Luna for Free and Go users starting the next day. explainx.ai answers what actually changed, what the 68% fewer-errors claim measures, and where the model picker went.

Jul 31, 2026

OpenAI Cuts GPT-5.6 Luna Price 80%, Terra 20% (July 2026)

OpenAI dropped GPT-5.6 Luna pricing 80% and Terra 20%, and shipped a Fast mode for Sol that runs up to 2.5x quicker at double the rate. The cuts apply automatically in Codex and ChatGPT Work usage accounting — here's what changed, why, and how Luna compares on cost per task against Claude and Gemini.

Jul 13, 2026

OpenAI Audits SWE-Bench Pro: ~30% of Tasks Broken — Retracts Recommendation

Frontier scores hit 80.3% on SWE-Bench Pro — then OpenAI flagged 27–34% broken tasks. Overly strict hidden tests, misleading prompts, verifier flaws. explainx.ai updates the benchmark trust stack after OpenAI retracts its Pro recommendation.