explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What's confirmed: Gemini 3.8 Flash and 3.8 Flash Cyber
  • TL;DR
  • Our original coverage, before the official launch
  • What's actually confirmed, versus what's a leak
  • The Flash release cadence this would extend
  • What we actually know about the Opus 5 coding comparison
  • Practitioner math: why this story matters even before it's confirmed
  • Does it inherit agentic video understanding?
  • Honest limitations
  • Closing
  • Related on explainx.ai
← Back to blog

explainx / blog

Gemini 3.8 Flash Is Official: Benchmarks, Flash Cyber, and Pricing

Gemini 3.8 Flash, Google Gemini, Model Launches, Coding Benchmarks, Claude Opus 5, Cybersecurity AI

Gemini 3.8 Flash and 3.8 Flash Cyber are officially live. Confirmed DeepSWE, HLE-Verified, and CyberGym benchmarks, the new Fairwind Program, and pricing.

Sep 2, 2026·14 min read·Yash Thakker
add explainx.ai
go deep
Gemini 3.8 Flash Is Official: Benchmarks, Flash Cyber, and Pricing

Update — September 24, 2026: New coverage — Gemini 3.8 Flash TTS.

Update — September 3, 2026: Google has officially confirmed Gemini 3.8 Flash and a new cybersecurity-focused Gemini 3.8 Flash Cyber, in a joint launch post from Tulsee Doshi (Senior Director of Product Management) and Raluca Ada Popa (Gemini Security Lead, Google DeepMind). The confirmed details are below. Our original leak-tracking coverage — published hours before the official announcement — is preserved further down this page as a record of what was known versus rumored at the time.

What's confirmed: Gemini 3.8 Flash and 3.8 Flash Cyber

Gemini 3.8 Flash is Google's third Flash-tier release in six weeks, arriving three weeks after Gemini 3.7 Flash. Google frames the gains as the model "working harder" on complex tasks — more reasoning steps, more iterative tool calls — rather than a new base architecture, and pricing stays at the same $0.75 per million input tokens / $3.75 per million output tokens introductory rate as 3.7 Flash, with a 1M token context window.

Confirmed benchmarks (Google's own reporting):

table · 3 cols
BenchmarkWhat it measuresGemini 3.8 Flash result
DeepSWE v1.1Long-horizon software engineeringOutperforms larger frontier models, per Google
HLE-VerifiedMulti-step reasoning (STEM, humanities, professional)54.9%
Vals Finance Agent V2Financial analysis and reportingImproves over 3.7 Flash and other frontier models
Harvey's Legal Agent BenchmarkLegal agent tasksImproves over 3.7 Flash and other frontier models

Google did not publish a same-benchmark, head-to-head score against Claude Opus 5 — the DeepSWE v1.1 result is compared against unnamed "larger frontier models," not a specific competitor table. That's the same caveat our original coverage below flags about the pre-launch leak claims: a vendor's own aggregate benchmark table is not the same as an independently run, apples-to-apples comparison.

Gemini 3.8 Flash Cyber, the cybersecurity-specialized sibling, prioritizes defensive capability — vulnerability discovery and automated patching — over offensive tooling. It scores frontier-level on CyberGym (autonomous vulnerability discovery) and lands on the Pareto frontier of CWE-Bench, run by Collinear, with 47.2% pass@1 on patching versus 47.8% for the leading frontier model, at meaningfully lower cost. Google cites real-world deployments: the Chrome Security team found it produced 2.6x more correct patches than the best commercial models tested; Wiz reported 7.5–9.7% higher recall on its internal penetration-testing benchmark at 2.3–5.2x lower cost; and Google's own Cloud Vulnerability Research team used it to find a critical vulnerability in under two hours.

Unlike the standard 3.8 Flash, Flash Cyber isn't generally available — it ships through a new Fairwind Program that gives prioritized access to trusted government authorities, critical infrastructure operators, and software maintainers, with a more permissive safety configuration than the general-release model reflecting its narrower, vetted audience.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
Is Gemini 3.8 Flash officially released?Yes, confirmed September 2, 2026
Confirmed pricing$0.75/$3.75 per million input/output tokens, 1M token context — same as 3.7 Flash
Confirmed benchmarksDeepSWE v1.1 (beats larger frontier models per Google), HLE-Verified 54.9%
Gemini 3.8 Flash CyberCybersecurity variant; frontier CyberGym score, CWE-Bench Pareto-frontier patching, defensive focus
How to access Flash CyberNew Fairwind Program — prioritized access for trusted defenders, not general availability
Head-to-head vs. Claude Opus 5Not published by Google — the DeepSWE claim compares against unnamed "larger frontier models"
Release cadence3.6 Flash + 3.5 Flash-Lite (July) → 3.7 Flash (Aug 13) → 3.8 Flash + Flash Cyber (Sept 2) — third release in six weeks
Real-world cyber results citedChrome Security: 2.6x more correct patches; Wiz: +7.5-9.7% recall at 2.3-5.2x lower cost; Google Cloud Vuln Research: critical bug found in under 2 hours

Our original coverage, before the official launch

The section below is preserved as originally published on September 2, 2026, tracking pre-launch leak reporting. It is kept for the record — everything above supersedes it with confirmed figures.

Headlines went out this week claiming Google is releasing Gemini 3.8 Flash on Wednesday, September 2, 2026, with a coding-benchmark edge over Claude Opus 5. We checked the primary sources before writing this — blog.google and ai.google.dev's model documentation — and neither shows a Gemini 3.8 Flash release as of this writing. The newest stable Flash model Google has actually documented is Gemini 3.7 Flash, which launched August 13, 2026.

That doesn't mean the story is baseless. A Wall Street Journal report (via Investing.com) and a cluster of leak sites — androidheadlines.com, nokiapoweruser.com, cryptobriefing.com, techgenyz.com, and shattered.io — all describe the same unreleased model, internally codenamed "skimaki," tested throughout August on Google's internal "Jetski" developer platform. This piece separates what's actually confirmed from what's still second-hand reporting, and walks through why the underlying question — can a cheap Flash-tier model really compete with an Opus-tier model on coding — matters for cost-conscious teams regardless of how Wednesday's launch actually shakes out.

Leak-era TL;DR (superseded, kept for the record)

table · 2 cols
QuestionAnswer at the time
Has Google confirmed Gemini 3.8 Flash?No. Not on blog.google or ai.google.dev as of that writing.
Where does the "Wednesday" claim come from?A WSJ report (via Investing.com) plus multiple leak/rumor sites — not a Google announcement
Internal codename"skimaki," reportedly tested on Google's internal "Jetski" developer platform through August 2026
The Opus 5 coding claimReported internal Google engineer preference in head-to-head Jetski testing — not a published benchmark score
Confirmed prior model: Gemini 3.7 Flash pricing$0.75/$3.75 per million input/output tokens through Dec 31, 2026; 1M token context window
Confirmed prior model: Claude Opus 5 pricing$5/$25 per million input/output tokens, launched July 24, 2026
Release cadence3.6 Flash + 3.5 Flash-Lite (July) → 3.7 Flash (Aug 13) → 3.8 Flash rumored (Sept 2) — roughly every 3 weeks
Agentic video understanding on 3.8 Flash?Not listed — Google's Sept 1 announcement names 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite only
Our read at the timeTreat the coding-edge claim as anticipated, not verified, until Google publishes a model card and benchmark table

What's actually confirmed, versus what's a leak

This is worth stating plainly before anything else, because the headline compresses "reportedly launching" into "released," and those are different claims with different reliability.

Confirmed, independently:

  • Gemini 3.7 Flash is real and current — launched August 13, 2026, at $0.75/$3.75 per million input/output tokens with a 1M-token context window, per Google's own launch post.
  • Google's agentic video understanding feature shipped September 1, 2026, on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
  • Claude Opus 5 launched July 24, 2026, at $5/$25 per million input/output tokens, per Ramp's enterprise spend data.
  • As of this writing, ai.google.dev's model list tops out at Gemini 3.7 Flash — no 3.8 entry.

Not confirmed — reported by leak sites and a single WSJ story:

  • A specific "Wednesday" launch date for Gemini 3.8 Flash
  • The internal codename "skimaki" and the "Jetski" testing platform
  • Any numeric benchmark score for Gemini 3.8 Flash against Claude Opus 5, on any named benchmark
  • Final pricing or context window for the new model

The distinction matters because "Google engineers preferred it in internal testing" — the actual substance behind the coding-edge headline, per the WSJ-sourced reporting — is a very different claim from "Gemini 3.8 Flash scored X% on SWE-bench versus Opus 5's Y%." The former is subjective, internal, and unverifiable from outside Google. The latter is the kind of claim that needs an actual published number before it should change anyone's model-selection decision. Right now, only the first kind of claim exists in public reporting.

The Flash release cadence this would extend

Whatever ships on Wednesday, it lands inside a genuinely fast release pattern Google has kept up since mid-2026. This is useful context on its own, separate from whether 3.8 Flash specifically is real yet.

table · 3 cols
ReleaseDateWhat changed
Gemini 3.6 Flash + 3.5 Flash-LiteJuly 2026Cyber-focused launch, price cuts across the Flash tier
Gemini 3.7 FlashAugust 13, 2026"Most intelligent workhorse yet for coding and agents," 1M context, $0.75/$3.75 pricing
Gemini 3.8 Flash (reported)September 2, 2026Reported focus: less verbose output, sharper multi-step agentic workflows, faster code generation, fewer long-chat hallucinations

Three weeks separated 3.6 Flash from 3.7 Flash. If the reporting holds, roughly three weeks separate 3.7 Flash from 3.8 Flash — consistent with CEO Sundar Pichai's stated target of close to a monthly cadence for Flash-tier models. Early tester feedback cited in the reporting frames 3.8 Flash as a refinement rather than a leap: cutting verbosity (a persistent complaint about earlier Flash models), tightening agentic tool-call sequencing, and reducing hallucinations in extended conversations, rather than a wholesale architecture change.

This matters for the benchmark claim specifically. A three-week iteration cycle is short enough that a real coding-benchmark jump over the prior model is plausible — Google's own 3.7 Flash showcase demonstrated genuine one-shot capability gains in that window — but it's also short enough that "beats Opus 5 on coding" deserves scrutiny rather than automatic belief, especially secondhand.

What we actually know about the Opus 5 coding comparison

Coding benchmark comparisons between Gemini Flash models and Claude Opus 5 already exist for the confirmed prior release, and they're a useful baseline for judging the 3.8 Flash claim once real numbers appear.

For Gemini 3.7 Flash specifically:

table · 4 cols
BenchmarkGemini 3.7 FlashClaude Opus 5Note
DeepSWE v1.165.3%Not directly reported on this benchmarkGoogle's own harness
SWE-bench ProNot directly reported79.2%Anthropic's own harness
Artificial Analysis Intelligence Index5663Aggregate reasoning score, not coding-specific

The honest read, per third-party comparisons of the confirmed models: Opus 5 leads on hard, architectural coding and aggregate reasoning; Gemini 3.7 Flash wins on cost per completed task, raw output speed, and scoped implementation work. Those two benchmarks — DeepSWE v1.1 and SWE-bench Pro — run on different harnesses with different task pools, so a Flash-tier win on one and an Opus-tier win on the other isn't a contradiction, it's a reminder that "beats X on coding" always needs the "on which benchmark" qualifier attached.

If Gemini 3.8 Flash actually publishes a benchmark table on launch, the number worth checking first is whichever one is directly comparable to a number Anthropic has already published for Opus 5 — same benchmark, same harness. A same-benchmark win is meaningfully different from a same-day headline built on a different benchmark that simply looks favorable.

Practitioner math: why this story matters even before it's confirmed

Set aside the unconfirmed 3.8 Flash specifics for a moment, because the underlying economics are already real with confirmed pricing today, and they're the actual reason "cheap model beats expensive model on coding" is worth a cost-conscious team's attention.

Gemini 3.7 Flash's confirmed pricing is $0.75 per million input tokens and $3.75 per million output tokens. Claude Opus 5's confirmed pricing is $5 per million input tokens and $25 per million output tokens. That's roughly a 6.7x price gap on both input and output, today, independent of whatever ships Wednesday.

If a Flash-tier model — 3.7, 3.8, or whichever version is currently shipping when you read this — produces comparable or better results on your actual coding workload, that 6.7x gap compounds fast on agentic coding sessions specifically, because those sessions burn through many tool calls, long tool-output contexts, and repeated re-reads of a codebase. A single agentic coding run that costs $3 on Opus 5 pricing could plausibly cost under $0.50 on Flash-tier pricing for equivalent output — and that difference scales linearly with usage volume, which is exactly the shape of cost that adds up across a team running agents all day.

The practical takeaway for teams choosing between frontier and mid-tier models:

  1. Don't trust either vendor's aggregate benchmark score as your selection criterion. DeepSWE v1.1 and SWE-bench Pro measure different things on different harnesses. Run your own held-out task set from your actual repo.
  2. Test at the effort/reasoning level you'd actually run in production, not the maximum-effort configuration used in marketing benchmarks — cost gaps compound differently depending on how hard each model is pushed.
  3. Watch for whether a coding win generalizes past the specific benchmark used to claim it. A Flash-tier model winning on scoped, well-specified implementation tasks doesn't necessarily mean it wins on long-horizon, architecturally ambiguous work — the exact split the confirmed Opus 5 vs. Flash comparisons already show.
  4. Price the whole session, not the model card. Tool-call overhead, retries, and context re-reads on a botched agentic run can erase a per-token cost advantage quickly if the cheaper model needs more attempts to land the same result.

This is the same calculus explainx.ai has tracked across the broader mid-tier-versus-frontier debate — see the Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max comparison and why Fable 5 isn't the best default model for the pattern playing out across vendors, not just this one Gemini release.

Does it inherit agentic video understanding?

One detail worth flagging for anyone using Gemini for multimodal work, not just coding: Google's agentic video understanding feature — where the model decides what parts of a video to inspect instead of processing every frame, cutting video tokens up to 88% — shipped September 1, 2026, explicitly on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Gemini 3.8 Flash is not named, for the simple reason that it didn't exist in Google's documentation when that feature launched a day earlier.

Whether 3.8 Flash inherits the feature automatically on launch, needs a separate compatibility note, or ships with an even more efficient version of it is an open question worth checking directly against Google's updated model docs once — or if — a real announcement lands.

Honest limitations

  • Google's DeepSWE v1.1 claim compares 3.8 Flash against unnamed "larger frontier models," not a named, same-benchmark score for Claude Opus 5 specifically — treat cross-vendor coding superiority claims as directional until an independent benchmark run confirms them.
  • Gemini 3.8 Flash Cyber is not generally available; it ships only through the Fairwind Program to vetted government, infrastructure, and software-maintainer applicants, so the CyberGym and CWE-Bench numbers can't be reproduced by an ordinary developer today.
  • The real-world patching results (Chrome Security, Wiz, Google Cloud Vulnerability Research) are Google-reported case studies, not independently audited third-party benchmarks.
  • Whether 3.8 Flash inherits agentic video understanding, which launched a day earlier on 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, is not addressed in the 3.8 Flash launch post — check Google's model docs directly if that's part of your workload.

Closing

Gemini 3.8 Flash landed exactly on the cadence Google has kept since mid-2026 — a third Flash-tier release in six weeks — at unchanged pricing from 3.7 Flash and with genuinely published benchmark gains on DeepSWE v1.1 and HLE-Verified. The more interesting release is arguably Flash Cyber: a defensively-focused cybersecurity model with real deployment case studies (Chrome, Wiz, Google's own vulnerability research team) gated behind a new access program rather than shipped broadly, which says something about how seriously Google is treating dual-use risk in this specific capability area. For cost-conscious teams, the practitioner math from our original coverage still holds — benchmark either model against your own repo and task mix, not the vendor's aggregate score. Follow @explainx_ai for updates as more independent comparisons land.

Related on explainx.ai

  • Gemini for Windows: Alt+Space Desktop App (Sep 2026)
  • Muse Spark 1.3: Meta's Model Ties Opus 5 on Coding Benchmarks (same week)
  • Gemini 3.7 Flash Is Official: Confirmed Pricing and Benchmarks
  • Google AlphaEvolve: Gemini-Powered Evolutionary Code Optimization
  • Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Launch
  • Gemini 3.7 Flash Showcase: What Googlers Are One-Shotting
  • Gemini Agentic Video Understanding: How It Works
  • Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6 Sol Comparison
  • Opus 5 Overtakes Fable 5 in Enterprise Spend
  • Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max Comparison
  • Why Fable 5 Is Not the Best Default Model

Sources

  • Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (official launch post, September 2, 2026)
  • Google — Introducing Gemini 3.7 Flash
  • Google AI for Developers — Gemini models documentation
  • Investing.com — Google prepares Gemini 3.8 Flash to narrow AI coding gap, WSJ reports (pre-launch reporting)

This post was originally published September 2, 2026, tracking pre-launch leak reporting on Gemini 3.8 Flash, and was updated September 3, 2026, with confirmed details from Google's official launch post. The "Our original coverage" section is preserved for the record; the confirmed section above it supersedes it wherever the two disagree.

Spotted something out of date? Let us know.

People in this article

  • Sundar Pichai →CEO of Google and Alphabet
Explore people in AI →
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Aug 27, 2026

Gemini's Ox Alpha Timing Backlash: The Corrected Timeline

Three posts from Google Gemini team members got read as trolling a rival model's launch. The dates say otherwise: the posts are from August 22, 2026, and Z.ai did not reveal Ox Alpha as GLM-5.3-Flash until August 26. Logan Kilpatrick's public reply was substantially correct — and underneath the drama sits a real practitioner question about how free preview windows distort model evaluation.

Aug 13, 2026

Gemini 3.7 Flash Is Official: $0.75/1M Input, 1M Context, Confirmed Pricing

Google's official blog.google post confirmed Gemini 3.7 Flash on August 14, 2026 — $0.75/$3.75 per million input/output tokens through the end of 2026, a 1M-token context window, and availability across Antigravity, AI Studio, Android Studio, Gemini Enterprise, and Spark. The pricing leak turned out to be fully accurate; the Gemini 3.5 Pro and Sergey Brin RSI claims did not.

Sep 22, 2026

OpenAI Says Its Model Solved 100+ Open Math Problems, Forms Advisory Group

OpenAI published "Advisory Group on Mathematics and Artificial Intelligence" on September 21, 2026, confirming that an internal model it began training August 28 has now resolved more than 100 long-standing open problems across most areas of mathematics — not just the Navier-Stokes Millennium Prize problem announced two weeks earlier. The same post forms an independent advisory group hosted at the Institute for Advanced Study, a direct response to 25 Fields Medalists' "severe misalignment" declaration from ten days prior. Here's what's confirmed, what the group can and can't do, and why the timing matters.