explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — what people are asking
  • What Frozen v2 reportedly changes
  • The compute shortage behind the chip
  • How this connects to Gemini 4's pre-training run
  • Why the market reacted the way it did
  • The TPU cadence Frozen v2 would sit alongside
  • What this means for developers and enterprises
  • Related reading
← Back to blog

explainx / blog

Google's Frozen v2 Chip and the Start of Gemini 4 Pre-Training

Google is reportedly building Frozen v2, a fixed-function chip beyond TPUs, to cut Nvidia reliance and power efficient Gemini inference — while confirming Gemini 4's pre-training run has begun. explainx.ai maps the chip strategy.

Jul 21, 2026·9 min read·Yash Thakker
Google TPUGemini 4AI InfrastructureCustom SiliconGoogle Cloud
go deep
Google's Frozen v2 Chip and the Start of Gemini 4 Pre-Training

Google is reportedly building a chip that only works if future Gemini models are designed around it — and simultaneously announced it has started training Gemini 4. Two pieces of Google infrastructure news landed within a day of each other in July 2026: reports of "Frozen v2," a fixed-function AI accelerator distinct from Google's TPU line, and Google's own confirmation that Gemini 4 pre-training has begun. Read together, they describe a company trying to buy itself out of a compute shortage severe enough that Google Cloud reportedly agreed to pay SpaceX close to $1 billion a month for capacity.

Here's what Frozen v2 is reported to do, why it's a different bet than the TPU line, and how it connects to the Gemini 4 training run Google just confirmed.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — what people are asking

QuestionAnswer
What is Frozen v2?A specialized AI chip beyond TPUs, reportedly 6-10x more tokens/watt
How does it save power?Reduces data movement, simplifies decision paths — fixed-function design
Trade-off?Flexibility — future Gemini models must match its architecture
Deployment target?~2028, per reports circulating July 20-21, 2026
Why now?Internal compute shortages; Google Cloud reportedly paying SpaceX ~$1B/month
Goal?Cut Nvidia GPU reliance, speed real-time AI like voice applications
Gemini 4 status?Pre-training confirmed started, per Google's July 21 announcement
Market reaction?Alphabet shares rose ~1.5% on the reporting

What Frozen v2 reportedly changes

The core claim is efficiency through specialization: Frozen v2 processes six to ten times more tokens per watt than Google's current infrastructure by cutting the two biggest costs in AI accelerator design — data movement between memory and compute, and the decision logic needed to support many different model shapes.

That's a meaningfully different design philosophy than the TPU line Google showcased at Cloud Next 2026 with TPU-8. TPUs are still programmable accelerators — they run whatever model architecture Google's researchers design next, with software adapting to new attention mechanisms, mixture-of-experts routing, or multimodal fusion approaches as Gemini evolves. Frozen v2 gives that flexibility up. The efficiency gain comes specifically from locking in architectural assumptions ahead of time, which means Google's model team would need to design future Gemini generations to fit the chip, not the other way around.

That's an unusual sequencing for a company that ships a new Gemini generation every few months. It only makes sense if Google is willing to slow model architecture experimentation in exchange for inference costs low enough to run agentic and real-time workloads — like voice applications — at a scale GPUs and even general-purpose TPUs can't hit economically.

Why "fixed-function" is the trade Google is reportedly willing to make

Every accelerator generation faces the same tension: general-purpose flexibility costs power and silicon area that a narrower, purpose-built design doesn't need. Nvidia's GPUs and Google's own TPUs have historically leaned toward flexibility because model architectures were still evolving fast enough that betting on a fixed shape risked stranding the hardware investment. A 6-10x tokens-per-watt jump is the kind of number that only shows up when a chip stops paying that flexibility tax — which is consistent with what's reported about Frozen v2's approach of reducing "data movement" and simplifying "AI decisions" at the hardware level.

The catch is co-design risk. If Frozen v2 assumes a specific attention mechanism or routing scheme and Gemini's research team needs to deviate from it for the next capability jump, Google either delays the model or under-uses the chip. That's the same kind of hardware-software coupling Apple and Tesla have bet on with their own custom silicon — accepting less flexibility for a step-change in efficiency on the workload that matters most.

The compute shortage behind the chip

None of this is happening in a vacuum. Google Cloud has reportedly prioritized internal Gemini compute needs over some external commitments, and agreed to pay SpaceX nearly $1 billion a month for additional capacity — a striking figure that signals just how constrained Google's own data center buildout has become relative to demand. That shortage is the backdrop for both the Frozen v2 reporting and the broader AI chip and memory pricing pressure explainx.ai has tracked through 2026, including the DRAM and HBM price surge hitting the entire industry.

Cutting Nvidia reliance isn't just a cost play — it's a supply security play. Every GPU Google doesn't need to buy is capacity it doesn't have to compete for against OpenAI, Anthropic, Meta, and every other frontier lab bidding for the same limited Nvidia allocation. A chip that's 6-10x more efficient per watt for Google's own inference workloads directly reduces that competition, even if it can never be sold externally the way TPUs are offered through Google Cloud.

How this connects to Gemini 4's pre-training run

The timing lines up with Google's own July 21, 2026 announcement, made alongside the Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber launch covered on explainx.ai's Gemini 3.6 Flash breakdown. Logan Kilpatrick's statement was direct: "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress."

Hacker News commenters connected the dots quickly. One reader speculated Google "waited for the new TPU generation to train a larger base model," while another suggested the six-month gap since Gemini 3's base pre-training run — without a competitive Pro-tier release in between — points to Google holding back its bigger training investment until infrastructure caught up. With Frozen v2's reported 2028 deployment target sitting well past Gemini 4's current training window, the likelier read is that Gemini 4 is training on existing or next-generation TPU infrastructure now, with Frozen v2 positioned for whatever model generation comes after — possibly Gemini 5 or a dedicated real-time/voice-optimized variant.

Why the market reacted the way it did

Alphabet shares reportedly rose around 1.5% on the Frozen v2 reporting — a modest but notable move for infrastructure news that hasn't been officially confirmed by Google. That reaction says more about investor anxiety over AI compute costs than about Frozen v2 itself. Every hyperscaler's margin story in 2026 has been shaped by the same question: how much of AI revenue gets eaten by Nvidia's take rate versus how much stays in-house through custom silicon. Amazon has Trainium, Microsoft has Maia, Meta has MTIA, and Google has had the longest-running custom silicon program of the group with TPUs stretching back nearly a decade. A report suggesting Google is willing to go even further — trading flexibility for a 6-10x efficiency jump — reads to investors as evidence Google intends to keep widening that in-house cost advantage rather than ceding more of it to Nvidia's supply constraints and pricing power.

That framing matters because it's easy to read "fixed-function chip" as a narrow technical footnote. In practice, it's a capital-allocation signal. Every dollar Google doesn't spend on Nvidia GPUs for internal inference is a dollar it either keeps as margin or redirects into training compute for the next Gemini generation — which is exactly the tradeoff visible in the same week's announcement that Gemini 4 pre-training has begun. Investors reading the two stories together are pricing in a company trying to compound efficiency gains on both sides of the ledger: cheaper inference silicon funding a bigger training run.

The TPU cadence Frozen v2 would sit alongside

Google's chip roadmap hasn't stood still while this reporting circulated. The Cloud Next 2026 TPU-8 announcement detailed a fleet built specifically to serve Gemini API traffic at scale, with cost and power-draw estimates already running into the billions of dollars annually and tens of megawatts of continuous draw. That existing TPU line is still the programmable, general-purpose workhorse — it has to flex across whatever architecture Gemini's research team ships next, from Flash-tier models to whatever Gemini 4 turns out to require.

Frozen v2, if the reporting holds up, would not replace that fleet. It would sit alongside it as a narrower, higher-efficiency option for the subset of inference workloads stable enough to justify hard-coding architectural assumptions into silicon — likely high-volume, latency-sensitive traffic like voice applications and agentic tool-calling loops, rather than novel research workloads still being iterated on. That's consistent with Google's own framing of the chip as targeting "real-time AI like voice applications" specifically, rather than positioning it as a blanket TPU replacement.

What this means for developers and enterprises

If you're...What to watch
Building on Gemini APIInference pricing could drop meaningfully once Frozen-class chips deploy (~2028) — but not before
Evaluating Google Cloud vs. other providersCompute shortage reporting suggests Google prioritizes internal Gemini workloads first
Comparing chip strategiesGoogle is betting narrower-and-more-efficient; Nvidia and most competitors still bet broader-and-flexible
Tracking Gemini 4No public timeline yet — pre-training just started, GA likely well into 2027 at the earliest

For teams building agentic workflows today, none of this changes anything immediately — Frozen v2 is years from deployment, and Gemini 4 pre-training runs typically take months before even reaching internal eval stages, a process explainx.ai covered in depth in the AI benchmarks complete guide. What it does signal is that Google's infrastructure roadmap and its model roadmap are now explicitly coupled — a shift worth tracking alongside Google's other 2026 chip and cloud moves and the ongoing AI bubble debate about whether this level of infrastructure spend is sustainable.

Related reading

  • GigaToken — Rust tokenizer claiming ~1000x faster, aimed at pretraining pipelines
  • Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber launch
  • Gemini 3.5 Pro benchmark leak — July 13, 2026
  • Google Cloud Next 2026: TPU-8 and Gemini Enterprise Agent Platform
  • Stanford memory prices: DRAM, HBM, NAND history 2026
  • Mobile DRAM price surge and AI smartphone shortage
  • AI bubble 2026: a reality check
  • AI climate change, energy, and sustainability

Details on Frozen v2 are based on reporting and social media discussion circulating July 20-21, 2026; Google has not published an official specification or roadmap for the chip as of this post's publication. Gemini 4 pre-training status reflects Google's own July 21, 2026 announcement. Chip and training timelines are subject to change — check Google's official developer and cloud channels for updates.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 13, 2026

The Next AI Bottleneck Is Pumps, Coolant, and Gas

A Polymarket post on August 12, 2026 said capital is rotating into vacuum pumps, cooling systems, and specialty gases for chip fabs and data centers. explainx.ai skips the ticker hunt and maps the physical bottleneck — and what it does to GPU availability and what you pay to run agents.

Aug 13, 2026

Anthropic Reportedly in Talks to Buy Decart for $6 Billion

Bloomberg reported on August 13, 2026 that Anthropic is in early-stage talks to acquire Israeli AI infrastructure and video-generation startup Decart for approximately $6 billion. Haaretz, Calcalist, i24News, and Yahoo Finance have corroborated the report, but nothing is signed and talks could still fall through. Here's what a deal this size would mean for Claude's compute capacity and inference roadmap if it closes.

Aug 12, 2026

Nvidia's $500B Plan to Make GPUs an Asset Class — and Why Its CDS Doubled

Jensen Huang wants GPU capacity treated like a toll road: an investable asset you borrow against. Six firms — Apollo, Blackstone, BlackRock, Brookfield, Goldman, KKR — signed on for $500 billion. The credit market reacted by pushing Nvidia's default swaps to a record. explainx.ai on what's actually being built and what the spread is pricing.