Google shipped Gemini 3.7 Flash on August 14, 2026 — just three weeks after Gemini 3.6 Flash — and called it, in Sundar Pichai's words, "a workhorse for performance at great value." Per Google's official announcement (Tulsee Doshi, blog.google, August 13, 2026) and its model card, the price is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50/$7.50 on January 1, 2027 — with a 1 million token context window and multimodal input, live day one in the Gemini API, AI Studio, and Antigravity. Google simultaneously cut Gemini 3.6 Flash's price to match ($0.75/$3.75) as of August 13, so the two Flash generations run at identical intro pricing right now. explainx.ai flagged this pricing the day before as an unverified X leak — the rumored input number turned out correct, though the leak never mentioned the output price, the price step-up in 2027, or the benchmark charts Google actually shipped with.
Those charts are the interesting part. Google's own launch data puts Gemini 3.7 Flash ahead of Claude Sonnet 5 and GPT-5.6 Terra on three of four published benchmarks — but the charts don't include Grok 4.6 at all, which fresh off its own August 12 launch is very much part of the conversation people are having about this tier of model. This piece lays out Google's numbers exactly as published, cross-references Grok 4.6 honestly from separate sources where it exists, and is explicit about the one place two different benchmarks share a name but not a scale — the same discipline explainx.ai applied comparing Fable 5, Grok 4.6, GPT-5.6 Sol, and Qwen3.8-Max and Sonnet 5 against GPT-5.6 Luna Max.
TL;DR
| Question | Direct answer |
|---|---|
| What did Google actually ship? | Gemini 3.7 Flash — $0.75/M input tokens through end of 2026, 1M context, multimodal, live in Gemini API / AI Studio / Antigravity |
| Does it beat Sonnet 5? | On Google's own chart, yes on 3 of 4 benchmarks — AutomationBench, Code Arena, and a near-tie on FrontierCode |
| Does it beat GPT-5.6 Terra? | Mixed — ahead on AutomationBench and Code Arena, roughly tied on FrontierCode, clearly behind on DeepSWE V1.1 |
| Where's Grok 4.6 on these charts? | Nowhere. Google's four benchmark charts don't include it at all — say so, don't guess |
| Can we compare Grok 4.6 anyway? | Only by cross-referencing separate sources: Grok 4.6 High scored 1618 Elo on a different Code Arena WebDev leaderboard, not Google's 1588 |
| Are the two "Code Arena" numbers the same test? | No — different panels (Google DeepMind's own vs. LMArena/Arena.ai), different absolute scales. Don't average them |
| Is this an independent benchmark? | No — it's Google's own vendor-published launch chart, same caveat as any lab's first-party numbers |
| Who actually wins DeepSWE V1.1? | GPT-5.6 Terra, at 69.6% — the one row where Gemini 3.7 Flash trails clearly |
Gemini 3.7 Flash: what actually shipped
Google's August 13-14, 2026 announcement — via official posts from @Google, @OfficialLoganK (Logan Kilpatrick), @GoogleDeepMind, and @GoogleAIStudio — pitches Gemini 3.7 Flash as "the most intelligent workhorse model yet for coding and agents." The headline numbers:
| Spec | Gemini 3.7 Flash |
|---|---|
| Input price | $0.75 per million tokens, held through end of 2026 |
| Context window | 1 million tokens |
| Modality | Multimodal |
| Availability | Gemini API, AI Studio, Antigravity — live at launch |
| Prior generation | Gemini 3.6 Flash, launched July 21, 2026 at $1.50/M input — roughly 3 weeks earlier |
A 50% input-price cut three weeks after the last Flash release is an aggressive cadence even by Google's own recent pace — 3.5 Flash (June) → 3.5 Flash-Lite and 3.6 Flash (July 21) → 3.7 Flash (August 14) is four Flash-tier releases inside nine weeks. That's the same rapid-iteration pattern explainx.ai flagged as a signal Google is defending the high-volume, cost-sensitive segment against GLM and DeepSeek pricing, not chasing frontier-reasoning parity outright.
The benchmark data — and its honesty caveat, upfront
Every number in the four tables below comes from Google DeepMind's own launch charts, methodology published at deepmind.google/models/evals-methodology/gemini-3-7-flash. That means this is first-party, vendor-published data — the same category as any lab's own launch chart, including SpaceXAI's Grok 4.6 table or Anthropic's own Sonnet 5 comparisons. Treat it the way explainx.ai treats any self-reported benchmark: directionally informative, not an independent audit. Nobody ran all four models in one neutral harness for this specific chart set.
The second, equally important caveat: Grok 4.6 is not on any of these four charts. Google's comparison set is its own generations plus Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Where this piece brings in a Grok 4.6 number, it's flagged explicitly as cross-referenced from a separate source — not read off the same chart.
AutomationBench — enterprise workflow automation
| Model | Score |
|---|---|
| Gemini 3.7 Flash | 30.4% |
| GPT-5.6 Terra | 23.6% |
| Gemini 3.6 Flash | 17.0% |
| Claude Sonnet 5 | 10.7% |
| Grok 4.6 | Not tested / no published score on this chart |
This is the largest single gap in Google's chart set. Gemini 3.7 Flash nearly doubles Gemini 3.6 Flash's own prior score (17.0% → 30.4%) in three weeks, and posts nearly 3x Claude Sonnet 5's 10.7%. AutomationBench measures enterprise workflow automation specifically — multi-step business-process tasks, not general coding — so this gap says more about agentic task-completion on structured business work than it does about raw coding ability.
Code Arena — web development (Google's chart)
| Model | Elo |
|---|---|
| Gemini 3.7 Flash | 1588 |
| Claude Sonnet 5 | 1541 |
| Muse Spark 1.2 | 1535 |
| Gemini 3.6 Flash | 1538 |
| GPT-5.6 Terra | 1523 |
| Grok 4.6 | Not on this chart |
Read this table carefully — it is not the same leaderboard as the Grok 4.6 comparison piece. explainx.ai's Fable 5 vs Grok 4.6 comparison cites a different "Code Arena WebDev" leaderboard, run by LMArena/Arena.ai, with entirely different absolute numbers: Fable 5 at 1627, GPT-5.6 Sol xHigh at 1622, Grok 4.6 High at 1618, Grok 4.5 at 1553. Both charts use the name "Code Arena" for a web-development coding evaluation, and both are Elo-style scores — but they are two different panels, judged by different raters, on a different scale. Gemini 3.7 Flash's 1588 on Google's own chart and Grok 4.6's 1618 on LMArena's chart are not directly comparable numbers, even though they look like they belong on the same axis. This is exactly the trap explainx.ai's guide to reading AI benchmarks warns about: a shared benchmark name is not a shared benchmark.
DeepSWE V1.1 — long-horizon software engineering
| Model | Score |
|---|---|
| GPT-5.6 Terra | 69.6% |
| Gemini 3.7 Flash | 65.3% |
| Muse Spark 1.2 | 54.9% |
| Claude Sonnet 5 | 53.8% |
| Gemini 3.6 Flash | 48.6% |
| Grok 4.6 | Not on this chart |
This is the one row where Gemini 3.7 Flash does not lead — GPT-5.6 Terra wins DeepSWE V1.1 outright, by more than 4 points. It's worth naming plainly rather than burying: Google's own launch chart shows its newest Flash model losing a benchmark to a competitor, and that's a more credible chart for including it. For color, SpaceXAI's own Grok 4.6 launch table separately reported Grok 4.6 High at 65.9% and GPT-5.6 Sol Max at 73% on the same-named DeepSWE V1.1 — a different GPT-5.6 tier (Sol, not Terra) scored on a different vendor's chart, so treat that as directional confirmation that DeepSWE V1.1 favors GPT-5.6's family generally, not a number you can merge into this table.
FrontierCode 1.1 Main — production code quality
| Model | Score |
|---|---|
| Gemini 3.7 Flash | 43.6% |
| Claude Sonnet 5 | 42.7% |
| GPT-5.6 Terra | 41.3% |
| Gemini 3.6 Flash | 34.4% |
| Grok 4.6 | Not on this chart |
The closest race on the whole chart set — Gemini 3.7 Flash, Sonnet 5, and GPT-5.6 Terra sit inside a 2.3-point band. Call this a statistical three-way tie on production code quality rather than a clear win; a 0.9-point edge over Sonnet 5 is well inside the noise you'd expect from any single-run vendor chart.
Beyond the four headline benchmarks: where Gemini 3.7 Flash does not lead
The four tables above are the ones the task at hand centers on, but Google's own launch chart runs wider than that — and the fuller table undercuts any temptation to read this as a clean Gemini sweep. Same source (Google DeepMind's evals methodology page), same five-model set:
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash | Claude Sonnet 5 | GPT-5.6 Terra | Muse Spark 1.2 |
|---|---|---|---|---|---|
| Artificial Analysis Intelligence Index | 56 | 52 | 55 | 57 | 57 |
| Terminal-bench 2.1 (agentic terminal coding) | 85.8% | 78.0% | 80.4% | 87.4% | 82.9% |
| Terminal-bench 3.0 (general agent capabilities) | 14.9% | 5.4% | 14.6% | 20.8% | — |
| GDPVal-AA v2 Elo (knowledge work) | 1525 | 1422 | 1598 | 1578 | 1628 |
| Harvey LAB-AA (complex legal workflows) | 90.7% | 85.1% | 90.1% | 85.2% | — |
| OSWorld-2.0 (agentic computer use) | 38.1% | 33.8% | 39.6% | 50.2% | — |
| HLE-Verified (multidisciplinary expert reasoning) | 53.6% | 51.2% | 31.0% | 51.1% | — |
| LVBench (long video understanding) | 85.4% | 84.2% | 68.5% | 78.9% | — |
Four separate models lead at least one row here, which is the honest picture: GPT-5.6 Terra actually leads the widest set — the AA Intelligence Index (tied with Muse Spark 1.2 at 57), both Terminal-bench versions, and OSWorld-2.0 by a wide 12-point margin. Muse Spark 1.2 leads GDPVal-AA v2 outright at 1628 Elo. Claude Sonnet 5 beats Gemini 3.7 Flash on GDPVal-AA v2 (1598 vs 1525) and OSWorld-2.0 (39.6% vs 38.1%) — real counterpoints to the AutomationBench and Code Arena leads covered above, and worth knowing if your workload leans toward knowledge-work or computer-use agents rather than coding. Gemini 3.7 Flash's clearest genuine strengths in this fuller table are HLE-Verified and LVBench — the latter unsurprising given Gemini's video-native multimodal design, the former a real reasoning-benchmark win where Sonnet 5 notably lags at 31.0%.
Terminal-bench 3.0's absolute scores (5.4%-20.8%) are worth a separate note: this is a hard, newer suite where every model in the set scores under 21%. Don't read GPT-5.6 Terra's win there the same way you'd read a 70%+ benchmark lead — it's a lead within a range where all five models are still mostly failing the task.
Cross-referencing Grok 4.6 honestly
Since Google's charts leave Grok 4.6 out entirely, the only responsible way to place it in this four-way conversation is to pull real numbers from where Grok 4.6 actually has been tested — and say clearly that none of it is the same chart as the numbers above.
| Source | Metric | Grok 4.6 result |
|---|---|---|
| SpaceXAI's own Aug 12 launch table | DeepSWE V1.1 | 65.9% (Grok 4.6 High) |
| LMArena/Arena.ai Code Arena WebDev | Elo | 1618 (Grok 4.6 High) — different leaderboard than Google's Code Arena above |
| Product Compass Bug Hunt Bench v10 | Bugs fixed / cost / time | 42 of 105 bugs, $22.73, 34 minutes |
| RuneScape Bench | Composite (⟨ln⟩) | 6.28 — highest of that panel |
None of these four numbers were measured against Gemini 3.7 Flash, Sonnet 5, or GPT-5.6 Terra in the same run. The DeepSWE V1.1 row is the closest thing to a genuine cross-check — both SpaceXAI and Google DeepMind published scores under the same benchmark name — and even there, Grok 4.6's 65.9% sits between Gemini 3.7 Flash's 65.3% and GPT-5.6 Terra's 69.6% on Google's chart, which is suggestive but not proof, since it's two different labs running the same-named suite independently. AutomationBench and FrontierCode 1.1 Main simply have no published Grok 4.6 number anywhere at the time of writing — not a gap this piece is going to paper over with an estimate.
One independent anecdote worth naming, with the same single-run caveat explainx.ai applied to the banana-drawing and cost-comparison threads in the Fable-5-vs-Grok-4.6 piece: a Hacker News post from developer jjcm (662 points, 376 comments) ran an image-to-HTML fidelity test — reproduce a source image as working HTML/CSS — across Claude Opus 5, Gemini 3.7 Flash, and Grok 4.6. Opus 5 won on visual polish, but several commenters flagged it as notable that Gemini 3.7 Flash outperformed Grok 4.6 on the same task, given Grok 4.6's momentum off its own recent launch. Treat that as one prompt, one judge, no scoring rubric — color that complicates a simple "Grok 4.6 is the value pick" narrative, not a benchmark result to weigh against the tables above.
Pricing snapshot
| Model | List price (per million tokens) | Notes |
|---|---|---|
| Gemini 3.7 Flash | $0.75 input / $3.75 output, through Dec 31, 2026 | Rises to $1.50/$7.50 on Jan 1, 2027; Google cut Gemini 3.6 Flash to match ($0.75/$3.75) the same week |
| Gemini 3.6 Flash | $0.75 input / $3.75 output (price-matched, Aug 13, 2026) | Was $1.50/$7.50 at its own July 21 launch — this is a mid-cycle cut, not the original price |
| Claude Sonnet 5 | $2 input / $10 output | Anthropic's permanent pricing as of August 11, 2026 |
| GPT-5.6 Terra | $2 input (per Google's own comparison chart) | Sits between Luna (cheapest) and Sol (priciest) in OpenAI's own GPT-5.6 tier structure |
| Muse Spark 1.2 | $1.25 input / $4.25 output (per Google's own comparison chart) | Included in Google's launch table as a fourth competitor, not covered elsewhere on explainx.ai yet |
| Grok 4.6 | $2 input / $6 output | Same as Grok 4.5; fast variant is 2x ($4/$12) — not part of Google's own pricing chart |
Gemini 3.7 Flash's $0.75 input rate undercuts Sonnet 5's $2 by more than 60% and Grok 4.6's $2 by the same margin — on sticker price alone, it's the cheapest model in this comparison during the intro window, before even accounting for the AutomationBench and Code Arena leads on Google's own chart. That gap narrows on January 1, 2027, when Gemini 3.7 Flash's own price doubles to $1.50/$7.50. As explainx.ai found comparing Sonnet 5 against GPT-5.6 Luna Max, sticker price and cost-per-completed-task diverge once you factor in tokens burned per task — nobody has published a cost-per-task figure for Gemini 3.7 Flash yet, so treat the pricing table as list price, not a verified cheapest-per-task claim.
When each model actually makes sense
| Priority | Best pick on this data | Why |
|---|---|---|
| Cheap, high-volume agentic coding | Gemini 3.7 Flash | Leads AutomationBench and Code Arena on Google's own chart, at roughly a third of Sonnet 5's or Grok 4.6's input price |
| Long-horizon software engineering | GPT-5.6 Terra | Only model to clearly win DeepSWE V1.1 on Google's chart (69.6%) |
| Production code quality at the margin | Any of the top three | Gemini 3.7 Flash, Sonnet 5, and GPT-5.6 Terra sit within 2.3 points on FrontierCode 1.1 Main |
| Independently-verified long-horizon agent execution | Grok 4.6, with caveats | Highest RuneScape Bench composite and fast, cheap Bug Hunt Bench run — but zero overlap with Google's own chart set |
| Claude Code / Anthropic ecosystem lock-in | Claude Sonnet 5 | Still the only option here with native Claude Code OAuth subscription billing in CI |
Bottom line
No single model wins this comparison cleanly, and the honest version of that sentence has three layers. First, on the four core benchmarks this piece leads with, Gemini 3.7 Flash tops three — AutomationBench by a wide margin, Code Arena web development, and a near-tie on FrontierCode — while GPT-5.6 Terra wins the fourth outright (DeepSWE V1.1, 69.6% to Gemini's 65.3%). Second, once you widen to Google's fuller benchmark set, GPT-5.6 Terra actually leads the most rows overall — both Terminal-bench versions and OSWorld-2.0 by a wide margin — and Muse Spark 1.2 and Claude Sonnet 5 each post real wins of their own (GDPVal-AA v2, and Sonnet 5 edges Gemini on OSWorld-2.0 too). This is not a Gemini sweep by any honest reading of Google's own data. Third, and just as important: this is Google's own vendor-published data, and Grok 4.6 isn't on any of it. Every Grok 4.6 number in this piece came from a different source, measuring a similarly-named benchmark on a different scale — treat the "Code Arena" comparison especially carefully, since Google's 1588 and LMArena's 1618 look like they belong on the same axis and don't. Gemini 3.7 Flash's price is real and aggressive at $0.75/M input tokens through the end of 2026; its benchmark lead is real on several of Google's own rows and needs independent confirmation before it's treated as settled. Run your own workload before switching a default.
Related on explainx.ai
- GPT-5.6 Sol Ultrafast mode: 750 TPS on Cerebras, same-day answer to this launch's speed positioning — HN's jjcm ties the two launches together directly
- Gemini 3.7 Flash pricing leak: the rumor before the confirmed launch — what was speculative on August 13 that Google confirmed a day later
- Grok 4.6 launch: official evals, pricing, Cursor access — the source for every Grok 4.6 number cross-referenced above
- Fable 5 vs Grok 4.6 vs GPT-5.6 Sol vs Qwen3.8-Max: full comparison — the source of the other Code Arena WebDev leaderboard cited above
- Claude Sonnet 5 vs GPT-5.6 Luna Max: cost comparison
- Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: what actually changed
- How to read AI benchmarks
- AI benchmarks complete guide
- Anthropic makes Claude Sonnet 5 pricing permanent
Sources: Introducing Gemini 3.7 Flash, Tulsee Doshi, blog.google, August 13, 2026 · Gemini 3.7 Flash model card, Google DeepMind · evals methodology at deepmind.google/models/evals-methodology/gemini-3-7-flash · Official X posts from @Google, @OfficialLoganK, @GoogleDeepMind, @GoogleAIStudio, August 13-14, 2026 · SpaceXAI Grok 4.6 launch, August 12, 2026 · Product Compass Bug Hunt Bench v10 · LMArena/Arena.ai Code Arena WebDev leaderboard · Hacker News, jjcm image-to-HTML fidelity thread (662 points, 376 comments)
Benchmark figures and pricing reflect August 14, 2026. AutomationBench, Code Arena, DeepSWE V1.1, and FrontierCode 1.1 Main figures for Gemini 3.7 Flash, Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2 come from Google DeepMind's own first-party launch chart, not an independent audit. Grok 4.6 does not appear on that chart; every Grok 4.6 figure in this piece is cross-referenced from a separate, independently-sourced benchmark and should not be read as measured on the same scale. Re-run your own workload before switching a production default.
