Update — September 24, 2026: New coverage — Gemini 3.8 Flash TTS.
Update — September 3, 2026: Google has officially confirmed Gemini 3.8 Flash and a new cybersecurity-focused Gemini 3.8 Flash Cyber, in a joint launch post from Tulsee Doshi (Senior Director of Product Management) and Raluca Ada Popa (Gemini Security Lead, Google DeepMind). The confirmed details are below. Our original leak-tracking coverage — published hours before the official announcement — is preserved further down this page as a record of what was known versus rumored at the time.
What's confirmed: Gemini 3.8 Flash and 3.8 Flash Cyber
Gemini 3.8 Flash is Google's third Flash-tier release in six weeks, arriving three weeks after Gemini 3.7 Flash. Google frames the gains as the model "working harder" on complex tasks — more reasoning steps, more iterative tool calls — rather than a new base architecture, and pricing stays at the same $0.75 per million input tokens / $3.75 per million output tokens introductory rate as 3.7 Flash, with a 1M token context window.
Confirmed benchmarks (Google's own reporting):
| Benchmark | What it measures | Gemini 3.8 Flash result |
|---|---|---|
| DeepSWE v1.1 | Long-horizon software engineering | Outperforms larger frontier models, per Google |
| HLE-Verified | Multi-step reasoning (STEM, humanities, professional) | 54.9% |
| Vals Finance Agent V2 | Financial analysis and reporting | Improves over 3.7 Flash and other frontier models |
| Harvey's Legal Agent Benchmark | Legal agent tasks | Improves over 3.7 Flash and other frontier models |
Google did not publish a same-benchmark, head-to-head score against Claude Opus 5 — the DeepSWE v1.1 result is compared against unnamed "larger frontier models," not a specific competitor table. That's the same caveat our original coverage below flags about the pre-launch leak claims: a vendor's own aggregate benchmark table is not the same as an independently run, apples-to-apples comparison.
Gemini 3.8 Flash Cyber, the cybersecurity-specialized sibling, prioritizes defensive capability — vulnerability discovery and automated patching — over offensive tooling. It scores frontier-level on CyberGym (autonomous vulnerability discovery) and lands on the Pareto frontier of CWE-Bench, run by Collinear, with 47.2% pass@1 on patching versus 47.8% for the leading frontier model, at meaningfully lower cost. Google cites real-world deployments: the Chrome Security team found it produced 2.6x more correct patches than the best commercial models tested; Wiz reported 7.5–9.7% higher recall on its internal penetration-testing benchmark at 2.3–5.2x lower cost; and Google's own Cloud Vulnerability Research team used it to find a critical vulnerability in under two hours.
Unlike the standard 3.8 Flash, Flash Cyber isn't generally available — it ships through a new Fairwind Program that gives prioritized access to trusted government authorities, critical infrastructure operators, and software maintainers, with a more permissive safety configuration than the general-release model reflecting its narrower, vetted audience.
TL;DR
| Question | Answer |
|---|---|
| Is Gemini 3.8 Flash officially released? | Yes, confirmed September 2, 2026 |
| Confirmed pricing | $0.75/$3.75 per million input/output tokens, 1M token context — same as 3.7 Flash |
| Confirmed benchmarks | DeepSWE v1.1 (beats larger frontier models per Google), HLE-Verified 54.9% |
| Gemini 3.8 Flash Cyber | Cybersecurity variant; frontier CyberGym score, CWE-Bench Pareto-frontier patching, defensive focus |
| How to access Flash Cyber | New Fairwind Program — prioritized access for trusted defenders, not general availability |
| Head-to-head vs. Claude Opus 5 | Not published by Google — the DeepSWE claim compares against unnamed "larger frontier models" |
| Release cadence | 3.6 Flash + 3.5 Flash-Lite (July) → 3.7 Flash (Aug 13) → 3.8 Flash + Flash Cyber (Sept 2) — third release in six weeks |
| Real-world cyber results cited | Chrome Security: 2.6x more correct patches; Wiz: +7.5-9.7% recall at 2.3-5.2x lower cost; Google Cloud Vuln Research: critical bug found in under 2 hours |
Our original coverage, before the official launch
The section below is preserved as originally published on September 2, 2026, tracking pre-launch leak reporting. It is kept for the record — everything above supersedes it with confirmed figures.
Headlines went out this week claiming Google is releasing Gemini 3.8 Flash on Wednesday, September 2, 2026, with a coding-benchmark edge over Claude Opus 5. We checked the primary sources before writing this — blog.google and ai.google.dev's model documentation — and neither shows a Gemini 3.8 Flash release as of this writing. The newest stable Flash model Google has actually documented is Gemini 3.7 Flash, which launched August 13, 2026.
That doesn't mean the story is baseless. A Wall Street Journal report (via Investing.com) and a cluster of leak sites — androidheadlines.com, nokiapoweruser.com, cryptobriefing.com, techgenyz.com, and shattered.io — all describe the same unreleased model, internally codenamed "skimaki," tested throughout August on Google's internal "Jetski" developer platform. This piece separates what's actually confirmed from what's still second-hand reporting, and walks through why the underlying question — can a cheap Flash-tier model really compete with an Opus-tier model on coding — matters for cost-conscious teams regardless of how Wednesday's launch actually shakes out.
Leak-era TL;DR (superseded, kept for the record)
| Question | Answer at the time |
|---|---|
| Has Google confirmed Gemini 3.8 Flash? | No. Not on blog.google or ai.google.dev as of that writing. |
| Where does the "Wednesday" claim come from? | A WSJ report (via Investing.com) plus multiple leak/rumor sites — not a Google announcement |
| Internal codename | "skimaki," reportedly tested on Google's internal "Jetski" developer platform through August 2026 |
| The Opus 5 coding claim | Reported internal Google engineer preference in head-to-head Jetski testing — not a published benchmark score |
| Confirmed prior model: Gemini 3.7 Flash pricing | $0.75/$3.75 per million input/output tokens through Dec 31, 2026; 1M token context window |
| Confirmed prior model: Claude Opus 5 pricing | $5/$25 per million input/output tokens, launched July 24, 2026 |
| Release cadence | 3.6 Flash + 3.5 Flash-Lite (July) → 3.7 Flash (Aug 13) → 3.8 Flash rumored (Sept 2) — roughly every 3 weeks |
| Agentic video understanding on 3.8 Flash? | Not listed — Google's Sept 1 announcement names 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite only |
| Our read at the time | Treat the coding-edge claim as anticipated, not verified, until Google publishes a model card and benchmark table |
What's actually confirmed, versus what's a leak
This is worth stating plainly before anything else, because the headline compresses "reportedly launching" into "released," and those are different claims with different reliability.
Confirmed, independently:
- Gemini 3.7 Flash is real and current — launched August 13, 2026, at $0.75/$3.75 per million input/output tokens with a 1M-token context window, per Google's own launch post.
- Google's agentic video understanding feature shipped September 1, 2026, on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
- Claude Opus 5 launched July 24, 2026, at $5/$25 per million input/output tokens, per Ramp's enterprise spend data.
- As of this writing, ai.google.dev's model list tops out at Gemini 3.7 Flash — no 3.8 entry.
Not confirmed — reported by leak sites and a single WSJ story:
- A specific "Wednesday" launch date for Gemini 3.8 Flash
- The internal codename "skimaki" and the "Jetski" testing platform
- Any numeric benchmark score for Gemini 3.8 Flash against Claude Opus 5, on any named benchmark
- Final pricing or context window for the new model
The distinction matters because "Google engineers preferred it in internal testing" — the actual substance behind the coding-edge headline, per the WSJ-sourced reporting — is a very different claim from "Gemini 3.8 Flash scored X% on SWE-bench versus Opus 5's Y%." The former is subjective, internal, and unverifiable from outside Google. The latter is the kind of claim that needs an actual published number before it should change anyone's model-selection decision. Right now, only the first kind of claim exists in public reporting.
The Flash release cadence this would extend
Whatever ships on Wednesday, it lands inside a genuinely fast release pattern Google has kept up since mid-2026. This is useful context on its own, separate from whether 3.8 Flash specifically is real yet.
| Release | Date | What changed |
|---|---|---|
| Gemini 3.6 Flash + 3.5 Flash-Lite | July 2026 | Cyber-focused launch, price cuts across the Flash tier |
| Gemini 3.7 Flash | August 13, 2026 | "Most intelligent workhorse yet for coding and agents," 1M context, $0.75/$3.75 pricing |
| Gemini 3.8 Flash (reported) | September 2, 2026 | Reported focus: less verbose output, sharper multi-step agentic workflows, faster code generation, fewer long-chat hallucinations |
Three weeks separated 3.6 Flash from 3.7 Flash. If the reporting holds, roughly three weeks separate 3.7 Flash from 3.8 Flash — consistent with CEO Sundar Pichai's stated target of close to a monthly cadence for Flash-tier models. Early tester feedback cited in the reporting frames 3.8 Flash as a refinement rather than a leap: cutting verbosity (a persistent complaint about earlier Flash models), tightening agentic tool-call sequencing, and reducing hallucinations in extended conversations, rather than a wholesale architecture change.
This matters for the benchmark claim specifically. A three-week iteration cycle is short enough that a real coding-benchmark jump over the prior model is plausible — Google's own 3.7 Flash showcase demonstrated genuine one-shot capability gains in that window — but it's also short enough that "beats Opus 5 on coding" deserves scrutiny rather than automatic belief, especially secondhand.
What we actually know about the Opus 5 coding comparison
Coding benchmark comparisons between Gemini Flash models and Claude Opus 5 already exist for the confirmed prior release, and they're a useful baseline for judging the 3.8 Flash claim once real numbers appear.
For Gemini 3.7 Flash specifically:
| Benchmark | Gemini 3.7 Flash | Claude Opus 5 | Note |
|---|---|---|---|
| DeepSWE v1.1 | 65.3% | Not directly reported on this benchmark | Google's own harness |
| SWE-bench Pro | Not directly reported | 79.2% | Anthropic's own harness |
| Artificial Analysis Intelligence Index | 56 | 63 | Aggregate reasoning score, not coding-specific |
The honest read, per third-party comparisons of the confirmed models: Opus 5 leads on hard, architectural coding and aggregate reasoning; Gemini 3.7 Flash wins on cost per completed task, raw output speed, and scoped implementation work. Those two benchmarks — DeepSWE v1.1 and SWE-bench Pro — run on different harnesses with different task pools, so a Flash-tier win on one and an Opus-tier win on the other isn't a contradiction, it's a reminder that "beats X on coding" always needs the "on which benchmark" qualifier attached.
If Gemini 3.8 Flash actually publishes a benchmark table on launch, the number worth checking first is whichever one is directly comparable to a number Anthropic has already published for Opus 5 — same benchmark, same harness. A same-benchmark win is meaningfully different from a same-day headline built on a different benchmark that simply looks favorable.
Practitioner math: why this story matters even before it's confirmed
Set aside the unconfirmed 3.8 Flash specifics for a moment, because the underlying economics are already real with confirmed pricing today, and they're the actual reason "cheap model beats expensive model on coding" is worth a cost-conscious team's attention.
Gemini 3.7 Flash's confirmed pricing is $0.75 per million input tokens and $3.75 per million output tokens. Claude Opus 5's confirmed pricing is $5 per million input tokens and $25 per million output tokens. That's roughly a 6.7x price gap on both input and output, today, independent of whatever ships Wednesday.
If a Flash-tier model — 3.7, 3.8, or whichever version is currently shipping when you read this — produces comparable or better results on your actual coding workload, that 6.7x gap compounds fast on agentic coding sessions specifically, because those sessions burn through many tool calls, long tool-output contexts, and repeated re-reads of a codebase. A single agentic coding run that costs $3 on Opus 5 pricing could plausibly cost under $0.50 on Flash-tier pricing for equivalent output — and that difference scales linearly with usage volume, which is exactly the shape of cost that adds up across a team running agents all day.
The practical takeaway for teams choosing between frontier and mid-tier models:
- Don't trust either vendor's aggregate benchmark score as your selection criterion. DeepSWE v1.1 and SWE-bench Pro measure different things on different harnesses. Run your own held-out task set from your actual repo.
- Test at the effort/reasoning level you'd actually run in production, not the maximum-effort configuration used in marketing benchmarks — cost gaps compound differently depending on how hard each model is pushed.
- Watch for whether a coding win generalizes past the specific benchmark used to claim it. A Flash-tier model winning on scoped, well-specified implementation tasks doesn't necessarily mean it wins on long-horizon, architecturally ambiguous work — the exact split the confirmed Opus 5 vs. Flash comparisons already show.
- Price the whole session, not the model card. Tool-call overhead, retries, and context re-reads on a botched agentic run can erase a per-token cost advantage quickly if the cheaper model needs more attempts to land the same result.
This is the same calculus explainx.ai has tracked across the broader mid-tier-versus-frontier debate — see the Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max comparison and why Fable 5 isn't the best default model for the pattern playing out across vendors, not just this one Gemini release.
Does it inherit agentic video understanding?
One detail worth flagging for anyone using Gemini for multimodal work, not just coding: Google's agentic video understanding feature — where the model decides what parts of a video to inspect instead of processing every frame, cutting video tokens up to 88% — shipped September 1, 2026, explicitly on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Gemini 3.8 Flash is not named, for the simple reason that it didn't exist in Google's documentation when that feature launched a day earlier.
Whether 3.8 Flash inherits the feature automatically on launch, needs a separate compatibility note, or ships with an even more efficient version of it is an open question worth checking directly against Google's updated model docs once — or if — a real announcement lands.
Honest limitations
- Google's DeepSWE v1.1 claim compares 3.8 Flash against unnamed "larger frontier models," not a named, same-benchmark score for Claude Opus 5 specifically — treat cross-vendor coding superiority claims as directional until an independent benchmark run confirms them.
- Gemini 3.8 Flash Cyber is not generally available; it ships only through the Fairwind Program to vetted government, infrastructure, and software-maintainer applicants, so the CyberGym and CWE-Bench numbers can't be reproduced by an ordinary developer today.
- The real-world patching results (Chrome Security, Wiz, Google Cloud Vulnerability Research) are Google-reported case studies, not independently audited third-party benchmarks.
- Whether 3.8 Flash inherits agentic video understanding, which launched a day earlier on 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, is not addressed in the 3.8 Flash launch post — check Google's model docs directly if that's part of your workload.
Closing
Gemini 3.8 Flash landed exactly on the cadence Google has kept since mid-2026 — a third Flash-tier release in six weeks — at unchanged pricing from 3.7 Flash and with genuinely published benchmark gains on DeepSWE v1.1 and HLE-Verified. The more interesting release is arguably Flash Cyber: a defensively-focused cybersecurity model with real deployment case studies (Chrome, Wiz, Google's own vulnerability research team) gated behind a new access program rather than shipped broadly, which says something about how seriously Google is treating dual-use risk in this specific capability area. For cost-conscious teams, the practitioner math from our original coverage still holds — benchmark either model against your own repo and task mix, not the vendor's aggregate score. Follow @explainx_ai for updates as more independent comparisons land.
Related on explainx.ai
- Gemini for Windows: Alt+Space Desktop App (Sep 2026)
- Muse Spark 1.3: Meta's Model Ties Opus 5 on Coding Benchmarks (same week)
- Gemini 3.7 Flash Is Official: Confirmed Pricing and Benchmarks
- Google AlphaEvolve: Gemini-Powered Evolutionary Code Optimization
- Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Launch
- Gemini 3.7 Flash Showcase: What Googlers Are One-Shotting
- Gemini Agentic Video Understanding: How It Works
- Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 Sol Comparison
- Opus 5 Overtakes Fable 5 in Enterprise Spend
- Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max Comparison
- Why Fable 5 Is Not the Best Default Model
Sources
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (official launch post, September 2, 2026)
- Google — Introducing Gemini 3.7 Flash
- Google AI for Developers — Gemini models documentation
- Investing.com — Google prepares Gemini 3.8 Flash to narrow AI coding gap, WSJ reports (pre-launch reporting)
This post was originally published September 2, 2026, tracking pre-launch leak reporting on Gemini 3.8 Flash, and was updated September 3, 2026, with confirmed details from Google's official launch post. The "Our original coverage" section is preserved for the record; the confirmed section above it supersedes it wherever the two disagree.
