explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • First measured results (August 25, 2026)
  • What Hacker News is debating
  • Why OpenAI Built Its Own Silicon
  • Nine-Month Tape-Out: A Record in Advanced Semiconductors
  • Architecture: What "Blank Slate for LLM Inference" Means
  • What GPT-5.3-Codex-Spark Running in Lab Means
  • The Full-Stack Strategy: Products → Models → Infrastructure → Silicon
  • Partner Roles: Broadcom and Celestica
  • Gigawatt Scale: What That Actually Means
  • Implications for the AI Infrastructure Landscape
  • Technical Deep Dive: Why Inference Silicon Differs from Training Silicon
  • What This Means for AI Education and Skill Building
  • What to Watch For
← Back to blog

explainx / blog

OpenAI Jalapeño: First AI Chip Built from Scratch for LLM Inference, Co-Developed with Broadcom

OpenAI unveiled Jalapeño, its first custom Intelligence Processor, co-designed with Broadcom from blank slate to tape-out in nine months. Purpose-built for LLM inference, it runs GPT-5.3-Codex-Spark in lab testing, targets performance-per-watt well above current state-of-the-art, and will deploy at gigawatt scale with Microsoft and other partners beginning in 2026.

Jun 24, 2026·27 min read·Yash Thakker
OpenAIAI HardwareLLM InferenceBroadcomAI InfrastructureChip Design
go deep
OpenAI Jalapeño: First AI Chip Built from Scratch for LLM Inference, Co-Developed with Broadcom

Updated — August 26, 2026: OpenAI published the first measured InferenceX benchmark results for Jalapeño on August 25 — not just lab anecdotes. SemiAnalysis engineers verified runs in OpenAI's lab. A 322-point Hacker News thread on the SemiAnalysis write-up pushed practitioners to stress-test the claims. Jump to First measured results for charts and numbers, or What Hacker News is debating for the synthesis.

OpenAI and Broadcom (NASDAQ: AVGO) on June 24, 2026 unveiled Jalapeño, OpenAI's first custom AI accelerator—an Intelligence Processor designed from scratch for LLM inference, not adapted from a general-purpose GPU lineage. The chip was delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom President and CEO Hock Tan and President Charlie Kawwas. Engineering samples are already running GPT-5.3-Codex-Spark at production-target frequency and power in the lab.

Primary sources:

  • Jalapeño's first results (Aug 25, 2026) — measured InferenceX benchmarks, architecture updates, deployment timeline.
  • OpenAI and Broadcom unveil LLM-optimized inference chip — original June announcement, chip architecture rationale, partnership scope.
  • SemiAnalysis: OpenAI Jalapeño vs Blackwell — independent lab verification, TCO analysis, AgentX caveats.
XSource postOpen on X ↗
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


First measured results (August 25, 2026)

On August 25, 2026, OpenAI published first benchmark numbers for Jalapeño — the June announcement's "coming months" technical report arriving two months early. SemiAnalysis engineers visited OpenAI's lab and verified InferenceX runs; the accompanying SemiAnalysis deep-dive adds TCO and architecture context. The X announcement thread cleared roughly 1M views within a day.

This is the story that matters for builders: not whether OpenAI has a chip, but whether measured serving efficiency can change API token economics the way GPT-5.6 Sol's August price cut already did on the software side.

TL;DR: InferenceX headline numbers

table · 2 cols
QuestionAnswer
What benchmark?SemiAnalysis InferenceX — nominal 8k input / 1k output context, single-turn serving (not multi-turn AgentX)
Compared against what?GB200 STP (GPT-OSS 120B) and GB300 STP (DeepSeek R1 670B, Kimi K2.5 1T) — single-token prediction baselines on Nvidia's best published configs
Chip power?700W TDP rated; ≤550W sustained measured on tested workloads
Peak efficiency gain?~1.5–1.9× more AI work per watt at peak mixed throughput
Latency gain?~1.7–3.6× lower end-to-end latency at matched operating points
Interactive workloads?~2.1–4.1× higher performance for low-concurrency, latency-sensitive serving
Still using Nvidia?Yes — OpenAI says it continues deploying Nvidia for training and inference alongside Jalapeño
When deployed?Target end of 2026; Gen 2 in development, Gen 3 shaping

The headline ~1.9× peak throughput-per-kW comes from GPT-OSS 120B — the widest gap in the suite. The Pareto curve below spans all three open-weight models tested; the model-specific sections break out each comparison.

InferenceX Pareto frontier — Jalapeño (700W TDP) vs GB200 STP (1200W) on GPT-OSS 120B at 8k input / 1k output context; peak mixed TPS/kW ~85,448 vs ~44,960 (~1.9×)

InferenceX · 8k/1k · Jalapeño 700W vs GB200 1200W. Source: OpenAI Jalapeño first results, August 25, 2026.

GPT-OSS 120B vs GB200 STP

On the 120B open-weight model, Jalapeño's peak mixed tokens-per-second per kilowatt (TPS/kW) hit roughly 85,448 versus 44,960 on GB200 STP — about 1.9× at peak. At matched throughput, latency fell roughly 4.2×. At GB200's previous best time-between-tokens (TBT) of 535 tok/s/user, Jalapeño delivered up to ~53.7× higher throughput per watt.

GPT-OSS 120B Pareto frontier — Jalapeño vs GB200 STP on InferenceX 8k/1k; Jalapeño sustains ≤550W measured despite 700W TDP rating

Peak mixed TPS/kW: ~85,448 (Jalapeño) vs ~44,960 (GB200 STP). InferenceX · 8k/1k · 700W vs 1200W.

GPT-OSS 120B operating points — latency vs throughput at matched concurrency; ~4.2× lower latency at equivalent throughput

At GB200's prior-best TBT of 535 tok/s/user, Jalapeño reached ~53.7× higher TPS/kW. InferenceX · 8k/1k · Jalapeño 700W vs GB200 1200W.

DeepSeek R1 670B vs GB300 STP

On DeepSeek R1, Jalapeño delivered roughly 1.7× peak mixed TPS/kW, ~3.6× lower latency at matched throughput, and up to ~104.3× throughput per watt at GB300's prior-best TBT of 169 tok/s/user. SemiAnalysis noted Jalapeño hit 700+ tok/s/user at concurrency 1 on R1 using single-token prediction only — no speculative decoding, no prefill-decode disaggregation.

DeepSeek R1 670B Pareto frontier — Jalapeño (700W) vs GB300 STP on InferenceX 8k/1k; ~1.7× peak mixed TPS/kW, ~3.6× lower latency at matched throughput

At GB300's prior-best TBT of 169 tok/s/user, Jalapeño reached ~104.3× higher TPS/kW. InferenceX · 8k/1k · Jalapeño 700W vs GB300 STP.

Kimi K2.5 1T vs GB300 STP

On Moonshot's 1-trillion-parameter Kimi K2.5 — the largest public model in the test set — Jalapeño showed roughly 1.5× peak mixed TPS/kW, ~3.4× lower latency, and up to ~56.1× throughput per watt at GB300's prior-best TBT of 182 tok/s/user. OpenAI says internal frontier-model tests widened the advantage further, though those numbers are not public.

Kimi K2.5 1T operating points — trillion-parameter serving on InferenceX 8k/1k; ~1.5× peak TPS/kW vs GB300 STP, ~56.1× TPS/kW at prior-best TBT of 182 tok/s/user

InferenceX · 8k/1k · Jalapeño 700W vs GB300 STP. Largest open-weight model in the test suite.

Architecture choices behind the numbers

OpenAI's August post fills in design details the June teaser skipped:

  • Co-designed chip, memory, network, and rack — not a drop-in accelerator card
  • KV cache locality optimized for transformer serving patterns (see our AI chip architecture guide for why memory bandwidth dominates inference)
  • Balanced prefill and decode on homogeneous silicon — OpenAI explicitly rejected fixed prefill/decode pool splitting because workload mixes shift across model eras
  • AI-assisted chip design — nine months from tape-out to meaningful benchmarks; Codex + GPT-Astra brought three open-weight models to performance in two months on fresh silicon
  • HBM4 at 15.4 TB/s bandwidth per package — early adopter alongside Rubin-class parts
  • B0 stepping already in fab with ~25% perf/W improvement over A0 silicon tested here

Gen 2 is in development; Gen 3 is being shaped. OpenAI still deploys Nvidia for training and inference — Jalapeño is additive, not a wholesale swap.

What people are asking (and the honest caveats)

Is this a fair fight against Blackwell? SemiAnalysis argues the better comparison is Vera Rubin (also HBM4, similar deployment window) — not GB200/GB300 STP configs that already carry software maturity advantages. Even on that framing, Jalapeño's tokens per megawatt reportedly beat Rubin's July published MTP numbers — though Rubin lacks the same independent benchmark disclosure OpenAI granted here.

Why trust 8k/1k numbers for agent workloads? SemiAnalysis flags that InferenceX 8k/1k is single-turn — it does not stress routers, prefix-cache churn, KV offload, or multi-turn AgentX patterns. A chip that wins on 8k/1k can still lose on long-context agent loops. Treat these as necessary but not sufficient evidence.

No multi-token prediction? Jalapeño's wins came without MTP/speculative decoding. SemiAnalysis notes Rubin's published numbers used speculative decoding; if Jalapeño adds it later, cost-per-token could drop another 3–5× — but that's forward-looking, not measured here.

Custom silicon competition: AMD's Taalas acquisition attacks inference from the opposite direction (weights etched into silicon). Anthropic hiring Google's TPU founder Amir Salek signals another frontier lab building custom inference. Cerebras-powered GPT-5.6 Sol Ultrafast shows OpenAI already routes some models to non-Nvidia inference today.

What this means if you build on OpenAI APIs: Jalapeño deployment by end of 2026 is the timeline OpenAI reconfirmed. If production matches lab results, expect continued pressure on per-token pricing and lower latency floors for Codex agent loops — but none of that is guaranteed until silicon is serving your requests, not SemiAnalysis's benchmark harness.

What Hacker News is debating

The SemiAnalysis Jalapeño thread hit 322 points on Hacker News within a day of OpenAI's August 25 results — unusually high for hardware coverage. The commenters weren't disputing that custom inference silicon can win on efficiency; they were arguing about which tradeoffs matter, whether the benchmark setup is apples-to-apples, and whether savings reach API buyers. Below is explainx.ai's practitioner read of the eight themes that kept resurfacing.

table · 3 cols
HN themeWhat commenters arguedHonest limitation
Weight-baked ASICsAMD's Taalas acquisition and Etched-style chips bake a frozen model into silicon — faster than streaming weights from HBM, but locked to one model revision~2-year tapeout lag vs model churn; metal-mask ROM can shorten turnaround for minor updates
Complement, not replacementInference ASICs sit beside GPUs the way 3dfx sat beside CPUs — specialized accelerators, not wholesale swapsOpenAI still deploys Nvidia for training and inference; Jalapeño is additive at gigawatt scale
Speculative decoding gapJalapeño's published wins exclude MTP/speculative decoding; Rubin's headline numbers included it (~3–5× boost, workload-dependent)Headline perf/W gaps may narrow once both sides enable the same serving tricks
Token economicsCheaper per-token inference can raise total spend (Jevons paradox) if labs capture margin instead of passing savings throughCounterpoint: GPT-5.6 Sol's August API price cut shows OpenAI already passing some software-side savings to developers
Benchmark methodologySemiAnalysis's InferenceX suite is credible but not neutral — trust the benchmark, verify incentivesDylan Patel's Broadcom-adjacent disclosures drew minor thread friction; independent replication still matters
Production vs lab chartsVendor Pareto curves rarely survive six months of real traffic — routers, cache churn, and mixed workloads differ from 8k/1k single-turn harnessesOpenAI granted lab access; production AgentX numbers are still outstanding
Cerebras curveWafer-scale parts optimize a different latency/throughput point — CS-4 targets extreme decode speed; Jalapeño targets perf/W at datacenter scaleOpenAI already routes Ultrafast Sol through Cerebras — multiple inference backends, not one winner
Local vs datacenterFaster datacenter silicon pressures home-GPU economics; benchmarking on DeepSeek/Kimi signals open-weight models matter for fair comparisonsJalapeño is internal infrastructure — not a consumer product — but open-weight baselines help builders reproduce claims on other hardware

Weight-baked silicon vs flexible inference

Several HN commenters pointed at the opposite end of the custom-silicon spectrum: baking weights into the die rather than optimizing weight streaming. AMD's Taalas deal is the clearest public example — a demo stack running Llama 3.1 8B on etched weights via chatjimmy.ai, with AMD positioning it for agentic and edge workloads where model churn is slower. Etched and similar startups push the same idea from the startup side.

The practitioner tension is familiar from our AI chip architectures guide: programmability vs realized utilization. Jalapeño keeps flexibility — new kernels, new model shapes, Gen 2/3 roadmaps — at the cost of still moving weights through HBM. Taalas-class chips trade away model agility for an order-of-magnitude decode advantage on a frozen graph. Neither obviates the other; they target different deployment lifetimes.

Complementarity, not a GPU funeral

The 3dfx analogy recurred for a reason. Dedicated inference ASICs can outperform general-purpose GPUs on narrow serving metrics without replacing the CUDA training stack, the supply-chain relationships, or the burst-capacity GPUs provide during model launches. OpenAI's own August post reiterated that it continues deploying Nvidia alongside Jalapeño — consistent with HN skepticism that a single lab-verified chart implies an immediate break from Blackwell/Rubin procurement.

For builders, the actionable read is simpler: expect heterogeneous inference fleets — Jalapeño racks for high-volume ChatGPT/API paths, Nvidia clusters for training and overflow, Cerebras or other partners for latency-tier products — not a sudden "Jalapeño-only" API tier.

The speculative-decoding caveat (again, because HN won't let it go)

Commenter bjourne's point deserves its own bullet because it changes how you read the headline multiples. Jalapeño's August numbers are single-token prediction only. Nvidia's Rubin comparisons in SemiAnalysis's write-up included multi-token prediction / speculative decoding, which can add roughly 3–5× effective throughput on some decode-heavy workloads — but not uniformly across models or concurrency levels.

That does not invalidate Jalapeño's lead; it means the gap is workload- and feature-dependent. If OpenAI adds speculative decoding on Jalapeño before production, costs could fall further; if Rubin closes the perf/W gap without Jalapeño adding MTP, the narrative flips. Treat August's charts as a lower bound on Jalapeño's potential, not a ceiling on Rubin's.

GPT-OSS 120B operating points — single-token prediction baseline; Jalapeño wins without MTP or speculative decoding enabled

All published Jalapeño numbers exclude multi-token prediction. InferenceX · 8k/1k · 700W vs GB200 1200W.

Token economics: cheaper tokens, more tokens

HN's economics thread split two ways. One camp invoked Jevons paradox — efficiency gains often increase total consumption, so per-token cost falling does not guarantee your bill shrinks if agents and reasoning budgets expand. Another camp noted labs may capture margin rather than pass hardware savings to users, citing historical cloud pricing stickiness.

Both are fair priors. The counter-evidence on OpenAI's software side is real: the August GPT-5.6 Sol API price cut lowered input/output rates before Jalapeño ships. Hardware efficiency and software pricing move on different clocks. If you budget API spend, plan for lower unit costs and higher volume — the combination token economics posts have been flagging all year — rather than assuming either extreme.

Trust the benchmark, verify incentives

SemiAnalysis's InferenceX suite is the most transparent independent look Jalapeño has received — engineers in OpenAI's lab, open-weight models, published Pareto curves. That is materially stronger than a press-release FLOPS figure. HN also surfaced disclosure questions around Dylan Patel's Broadcom relationships in a side thread — not a reason to dismiss the numbers, but a reminder that independent replication (Artificial Analysis, MLPerf-style serving, or cloud A/B once silicon is live) still settles arguments.

explainx.ai's read: weight the methodology section (8k/1k single-turn, STP baselines, power caps) more heavily than the adjectives in the headline. See AI chip architectures for why memory-bandwidth-bound inference benchmarks are easy to miscompare across vendors.

GPT-OSS 120B Pareto frontier — the curve HN practitioners will stress-test once Jalapeño hits production traffic

Peak ~1.9× TPS/kW on GPT-OSS 120B. Lab-verified Pareto; production AgentX numbers still outstanding.

Production reality beats launch charts

Several practitioners noted the gap between vendor Pareto curves and six months of production traffic — mixed batch sizes, prefix-cache hits, routing layers, and agent loops that InferenceX's nominal 8k/1k single-turn profile does not simulate. Others highlighted scale and supply chain: OpenAI's gigawatt ambitions still run through Broadcom implementation, Celestica integration, and fab partners — custom silicon does not remove TSMC queue risk, it relocates it.

That is why SemiAnalysis itself flagged AgentX follow-ups. Jalapeño may win the chart and still face surprises when ChatGPT-scale routing lands on Gen 1 silicon. Watch for production telemetry, not just Hot Chips slides.

Where Cerebras and Trainium fit

HN commenters comparing Jalapeño to Cerebras CS-4 were not wrong to separate the products — wafer-scale engines optimize interactive decode at extreme tokens/sec, while Jalapeño optimizes perf/W across datacenter megawatts. OpenAI already demonstrated multi-vendor inference with GPT-5.6 Sol Ultrafast on Cerebras. AWS Trainium + Cerebras WSE combinations came up as another heterogeneous pattern: training/compute on one architecture, latency tier on another.

Open-weight models as the honest benchmark surface

Finally, HN noted that OpenAI benchmarked on GPT-OSS, DeepSeek R1, and Kimi K2.5 — open or semi-open weights builders can inspect — rather than opaque frontier checkpoints alone. That matters for reproducibility and for the local vs datacenter debate: when perf/W claims use public model graphs, independent teams can sanity-check serving behavior on GPUs they already own, even if they will never touch Jalapeño silicon.

DeepSeek R1 670B Pareto — open-weight 670B model; ~1.7× peak TPS/kW, 700+ tok/s/user at concurrency 1 without speculative decoding

InferenceX · 8k/1k · Jalapeño 700W vs GB300 STP.

Kimi K2.5 1T operating points — trillion-parameter open benchmark surface for independent perf/W verification

~1.5× peak TPS/kW; ~56.1× TPS/kW at GB300's prior-best TBT. InferenceX · 8k/1k.


Why OpenAI Built Its Own Silicon

The core argument Greg Brockman made in the announcement is simple: inference is where intelligence reaches people. Every ChatGPT response, every Codex task, every API call runs through inference hardware. If you own that hardware, you control cost, latency, and capacity in ways you cannot if you depend entirely on third-party silicon.

OpenAI has until now run on Nvidia GPUs—powerful, battle-tested, and expensive. GPUs were designed first for graphics, then adapted for matrix math, then further adapted for transformer models. Each layer of adaptation leaves efficiency on the table. Jalapeño starts from the premise: what if we designed a chip purely around the kernels, memory-movement patterns, and networking behaviors of modern LLMs?

Richard Ho, who leads OpenAI's hardware program, described it this way: "We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models."

The flywheel OpenAI is betting on:

  1. Custom silicon → better inference efficiency
  2. Better efficiency → cheaper and faster serving
  3. Cheaper serving → more capable products and lower API prices
  4. More usage → more revenue → reinvest in next-generation infrastructure
  5. Better infrastructure → next-generation models
  6. Better models → help design next chip (OpenAI models assisted Jalapeño's own design)

That last loop is the one Brockman highlighted explicitly: the same models served to users helped design the chip that will run future models.


Nine-Month Tape-Out: A Record in Advanced Semiconductors

Traditional ASIC development at advanced nodes (3nm, 5nm) typically takes 18–36 months from initial architecture to first silicon. Enterprise-class AI accelerators—TPU, Trainium, Gaudi—have generally required 2–3 years of dedicated engineering before reaching production.

OpenAI and Broadcom claim nine months from initial design to manufacturing tape-out. Three factors made this possible:

1. Deep Software-Hardware Co-Development

OpenAI writes and operates its own inference stack end-to-end: kernels, serving frameworks, batching logic, routing, and product APIs. That means the hardware team had direct access to the exact compute patterns to optimize for—no approximations from external benchmarks. The chip architecture was shaped by real production request traces from ChatGPT and the API.

2. Broadcom's Silicon Implementation Expertise

Broadcom has decades of experience implementing complex ASICs. Their Tomahawk networking silicon family powers some of the largest hyperscale data centers. Bringing that implementation discipline to OpenAI's architecture spec allowed parallel workstreams—architecture definition, physical design, verification, packaging—to proceed faster than a single-vendor effort could.

3. OpenAI Models Accelerating Chip Design

Perhaps the most forward-looking detail: OpenAI used its own models to accelerate "parts of the design and optimization process." This likely includes:

  • RTL generation: LLMs producing hardware description language (VHDL/SystemVerilog) for repetitive logic blocks
  • Verification test generation: Models writing exhaustive test cases for timing and functional correctness
  • Design-space exploration: Using AI to evaluate thousands of micro-architecture variants (cache sizes, pipeline depths, memory bandwidths) faster than human engineers could manually
  • Documentation and constraint generation: Auto-generating synthesis constraints and signoff checklists

This is significant because it suggests the chip-design-with-AI loop is already real and productive, not a roadmap slide.


Architecture: What "Blank Slate for LLM Inference" Means

OpenAI has not released a detailed datasheet yet (a full technical report is promised "in the coming months"), but the announcement describes the design philosophy in enough detail to understand the architectural bets:

Reducing Data Movement

Data movement is the primary bottleneck—and energy consumer—in LLM inference. Every token generation requires:

  1. Loading model weights from memory into compute units
  2. Computing attention over the KV cache (past context)
  3. Accumulating activations through MLP layers
  4. Sampling from the output distribution

In a general-purpose GPU, these memory loads traverse a generic memory hierarchy (registers → L1 → L2 → HBM) not tuned for transformer access patterns. Jalapeño, by contrast, was designed knowing exactly which data structures (weight matrices, KV tensors, activation buffers) get accessed in which order.

The announcement says the architecture "reduces data movement"—meaning the chip likely has:

  • Larger on-chip SRAM sized to hold key working sets
  • Custom memory controllers optimized for transformer weight shapes
  • Fused kernel support that keeps intermediate activations in registers rather than spilling to DRAM

Balancing Compute, Memory, and Networking

A common failure mode in AI accelerator design is roofline imbalance: more raw FLOPS than memory bandwidth can feed, leaving compute units idle waiting for data. Jalapeño was apparently designed to avoid this by co-optimizing all three dimensions against OpenAI's actual model shapes and serving patterns.

For an inference chip, the relevant ratios look different than for a training chip:

table · 3 cols
DimensionTraining priorityInference priority
Raw FLOPSMaximizeMatch to memory BW
Memory BWHighCritical (weight streaming)
On-chip SRAMModerateLarge (KV cache)
Network latencyTolerableLow (request tail latency)
Power per tokenSecondaryPrimary (cost structure)

OpenAI designs its models, kernels, and serving system—so the chip team could target the exact model widths and depths, attention head counts, and batch sizes that appear in production. This produces much better realized utilization vs. theoretical peak.

Networking via Broadcom Tomahawk

Broadcom's Tomahawk networking silicon is a hyperscale Ethernet switch series. At gigawatt-scale deployment, Jalapeño nodes will need to form large inference clusters for:

  • Tensor parallelism: Splitting a single large model across many chips
  • Pipeline parallelism: Distributing model layers across a deep chip pipeline
  • Disaggregated serving: Prefill nodes (processing input prompts) and decode nodes (generating tokens) on separate hardware

Tomahawk's high-radix, low-latency fabric is a natural fit—it's already deployed at the scale of thousands of servers in hyperscaler data centers. Integrating it into the Jalapeño platform from day one signals OpenAI is targeting not just single-chip performance but cluster efficiency at gigawatt scale.


What GPT-5.3-Codex-Spark Running in Lab Means

The fact that engineering samples are running GPT-5.3-Codex-Spark (a production model) at production-target frequency and power is a stronger signal than it might appear.

Most AI chip announcements show toy workloads or synthetic benchmarks in early silicon. Running a real production model means:

  • Firmware and kernel stack is functional end-to-end
  • Memory controllers and interconnect are delivering sufficient bandwidth for the model
  • Numerical correctness (FP8/FP16 outputs matching reference) is validated
  • Power and thermals are within target envelope at production frequency

"Production target frequency and power" means the chip isn't running slowly to avoid errors—it's at the operating point it will ship at. That's a meaningful milestone for a nine-month project.

OpenAI says "while we are still measuring final performance"—the final numbers haven't been locked, but engineering samples are functioning correctly at spec.


The Full-Stack Strategy: Products → Models → Infrastructure → Silicon

Jalapeño completes a vertical integration stack that previously stopped at software:

snippet
ChatGPT / Codex / API (Products)
    ↓
GPT-5.x / o-series (Models)
    ↓
Kernels / Serving / Routing (Software)
    ↓
Jalapeño + Tomahawk (Silicon)  ← new
    ↓
Data center (Microsoft, partners)

The analogy often cited is Apple's A-series chips. Apple designing both iOS and the A-chip allowed optimizations impossible when hardware and software came from separate companies. iPhone's battery life and performance-per-watt lead over Android for years was largely attributable to vertical integration.

For AI inference, the analogous opportunity is per-token cost and latency. If Jalapeño delivers the claimed performance-per-watt advantage over current state-of-the-art (a detailed number is forthcoming), OpenAI could:

  • Lower API prices without cutting margins
  • Increase context window affordably (longer KV caches need more memory bandwidth—custom silicon can deliver this more cheaply)
  • Prioritize capacity for products over external competitors who also use Nvidia
  • Reduce exposure to Nvidia supply constraints and pricing power

Partner Roles: Broadcom and Celestica

Broadcom

Broadcom's role goes beyond silicon implementation:

  • Chip implementation: Physical design, timing closure, DFT, packaging at advanced node
  • Networking silicon: Tomahawk switches for cluster-scale interconnect
  • Connectivity technologies: SerDes, optical interfaces, and system-level integration know-how

Broadcom has deep relationships with TSMC for advanced node manufacturing and experience managing the complexity of 5nm/3nm tape-outs. Their ecosystem of chiplet packaging and CoWoS-adjacent technologies may also be relevant for Jalapeño's memory subsystem.

Celestica

Celestica handles the mechanical and system layer:

  • Board design: PCB layout, signal integrity for high-speed memory interfaces
  • Rack integration: Power delivery, cooling, cabling at rack level
  • System production: Manufacturing at scale, quality control, supply chain management

Custom AI accelerators need custom boards—the connector layout, power delivery network, and thermal solution all differ from standard GPU server designs. Celestica's manufacturing scale enables the rapid ramp to gigawatt-scale deployment that Broadcom CEO Hock Tan referenced.


Gigawatt Scale: What That Actually Means

Hock Tan's quote—"deployment of gigawatt scale data centers with Microsoft and other partners beginning in 2026"—translates to enormous numbers.

Power math:

  • A modern AI accelerator consumes ~200–400W per chip
  • A 1 MW data center holds ~2,500–5,000 chips
  • A 1 GW deployment holds ~2.5–5 million chips

That's not one data center. It's a network of hyperscale facilities—consistent with Microsoft's widely reported $80B+ AI infrastructure investment commitment for 2025–2026, much of which is co-located with OpenAI.

Why inference, not training?

Training is a one-time (or periodic) cost. Inference is continuous—every ChatGPT user generates inference traffic every second they use the product. At OpenAI's scale (reported 400M+ weekly users), inference hardware needs to be massive to keep latency acceptable. Training can be done in bursts on rented GPU clusters; inference needs to be reliable, low-latency, and always available.


Implications for the AI Infrastructure Landscape

For OpenAI Users and Developers

If Jalapeño delivers on its performance-per-watt promise:

  • Faster responses: Lower latency for ChatGPT and Codex
  • Cheaper API: More efficient infrastructure can translate to lower per-token prices
  • Longer context: Memory-efficient inference enables larger context windows affordably
  • Better reliability: Owning the hardware stack reduces dependency on third-party supply chains

For Nvidia

Nvidia's data center GPU business has been the primary beneficiary of the AI boom. OpenAI is their largest reported customer. If Jalapeño can serve OpenAI's inference workloads more efficiently, it directly reduces future Nvidia GPU demand from OpenAI—at least for inference.

Training, however, remains a different story. Nvidia's H100/B200 architecture has a network effect around CUDA tooling, existing training frameworks, and researcher familiarity. OpenAI is unlikely to replace training hardware immediately even if Jalapeño succeeds at inference.

For Other AI Labs

The signal Jalapeño sends is clear: at sufficient scale, building custom inference silicon is worth it. Google has been doing this with TPUs since 2016. Amazon has Trainium (training) and Inferentia (inference). Meta has MTIA. OpenAI joining this group means the large labs are increasingly competing on infrastructure, not just on model quality.

Smaller labs without the revenue to justify ASIC development face a structural cost disadvantage as inference silicon becomes more efficient for each incumbent.

For Broadcom's Business

Broadcom CEO Hock Tan called this "a fundamental commitment to scaling the physical infrastructure required for the next decade of AI." The multi-generation chip platform agreement with OpenAI is a new, high-volume revenue stream diversified from Broadcom's networking switch business.

Broadcom also works with Google on TPU silicon and has similar relationships with Apple and other hyperscalers. Jalapeño adds OpenAI to their AI accelerator customer list.


Technical Deep Dive: Why Inference Silicon Differs from Training Silicon

Understanding why inference chips look different from training chips is useful context for anyone building on AI APIs.

Training vs. Inference: Different Bottlenecks

Training is compute-bound: the goal is to maximize FLOPS for gradient computation across billions of parameters.

Inference is memory-bandwidth-bound: the goal is to stream model weights (often tens to hundreds of gigabytes) through compute units as fast as possible to generate tokens one at a time.

A simple illustration for a 70B parameter model:

  • Weight size at FP8: 70B × 1 byte ≈ 70 GB
  • Generating 1 token requires reading all 70 GB (one forward pass)
  • Generating 100 tokens/second requires 7 TB/s of memory bandwidth

No current HBM technology delivers 7 TB/s on a single chip (H100 delivers ~3.35 TB/s). This forces batching: combining multiple users' requests so each weight load serves multiple tokens. The larger the batch, the more efficient the compute—but larger batches increase latency for individual users.

Jalapeño's design presumably addresses this tradeoff by:

  1. Maximizing HBM bandwidth per chip
  2. Maximizing on-chip SRAM to cache frequently-used weights and KV tensors
  3. Minimizing data movement by keeping activations on-chip across layers
  4. Optimizing batch scheduling to maximize realized throughput per watt

The KV Cache Problem

KV cache stores attention keys and values from earlier tokens in a conversation. It grows linearly with context length. For a 128K context window with a large model:

  • KV cache size per user ≈ 32 GB (approximate for a 70B model at 128K context)
  • For 1,000 concurrent users: 32 TB of KV cache storage

This means inference at scale requires either:

  • Very large on-chip or near-chip memory
  • Efficient KV cache compression (quantization, eviction policies)
  • Disaggregated memory tiers (fast NVMe or CXL-attached memory for older context)

Jalapeño's blank-slate design almost certainly includes specific hardware support for KV cache management—a feature absent in general-purpose GPUs.

Networking for Disaggregated Inference

Modern serving systems separate prefill (processing the input prompt—compute-heavy) from decode (generating output tokens one at a time—memory-bandwidth-heavy). Disaggregating these onto separate hardware pools improves utilization.

Jalapeño + Tomahawk enables fast interconnects between prefill nodes and decode nodes. This is similar to what Google described for TPU 8i's Boardfly topology and Collectives Acceleration Engine, suggesting the industry is converging on disaggregated inference as the serving paradigm for large-scale LLM products.


What This Means for AI Education and Skill Building

The Jalapeño announcement is easy to read as purely an infrastructure story. But it has clear implications for practitioners building on AI:

API Pricing Will Change

If Jalapeño delivers the efficiency gains OpenAI claims — the August InferenceX suite showed ~1.5–1.9× peak TPS/kW across open models at 8k/1k context — downstream API pricing will likely fall, continuing the trend of rapidly declining per-token costs. This matters for:

  • Agent pipelines with high token volumes (multi-turn reasoning, long document processing)
  • Real-time applications where latency is the current blocker
  • Startups that couldn't afford frontier-model API costs at their required scale

Budget your roadmap knowing inference costs are likely to fall over 2026–2027.

Infrastructure Ownership Is a Moat

For anyone building on OpenAI APIs: the company's vertical integration—from silicon to serving to product—creates a compounding advantage. Their cost structure for inference will improve faster than labs relying on third-party hardware. Build on APIs while understanding that the underlying economics will shift, and position your product on differentiation above the infrastructure layer.

The Full-Stack Flywheel Is Real

OpenAI using its own models to design the chip that will run future models is not a press release flourish—it's a concrete demonstration of the AI capability flywheel. Labs with deployed products and real revenue can invest those proceeds into infrastructure improvements that make future products better. This dynamic accelerates as the loop tightens.

Learn the Stack, Not Just the API

Understanding how inference hardware works—memory bandwidth, batching, KV cache—makes you a better AI engineer even if you never touch hardware. The constraints of inference silicon explain why context windows have pricing tiers, why latency varies with load, and why some model sizes are more cost-effective than others. These are skills that transfer across providers.

Courses and bootcamps at explainx.ai cover these fundamentals—not just "how to call the API" but "why the API behaves the way it does and how to build systems that use it efficiently."


What to Watch For

  • AgentX / multi-turn benchmarks: SemiAnalysis has not yet run full AgentX on Jalapeño — the August 25 results are InferenceX 8k/1k only. Long-context agent serving may tell a different story.
  • Rubin head-to-head: Fair comparison is Vera Rubin (HBM4, similar ship window), not GB200 STP alone. Watch whether Rubin closes the gap as Nvidia's software matures.
  • MTP / speculative decoding on Jalapeño: Current numbers exclude multi-token prediction. If OpenAI adds it, cost-per-token could fall further — but that's unmeasured today.
  • API price changes: Watch for pricing updates to GPT-5.3+ and Codex models in H2 2026 as Jalapeño comes online — see August Sol API cut for the software-side precedent.
  • Second generation: B0 stepping (~25% perf/W over A0) is already in fab. Gen 2 and Gen 3 are active programs.
  • Training silicon: Jalapeño is inference-only. Custom training silicon would be a deeper Nvidia break.
  • Competitive response: Google TPU 8t/8i, Amazon Trainium/Inferentia, AMD-Taalas, Anthropic's chip hiring — the AI chip landscape is moving fast.

Update — August 26, 2026: First measured InferenceX results are in — Jalapeño benchmark numbers show up to ~1.9× peak throughput per watt and ~4× lower latency vs GB200/GB300 on open models. Hacker News debate synthesis covers Taalas-style weight-baked chips, speculative-decoding caveats, and token economics.

Update — August 7, 2026: A more radical take on custom inference silicon just landed — AMD acquired Taalas, a startup that etches model weights directly into chip silicon rather than streaming them from HBM. Its test chip claims 16,960 tokens/second, 48x a typical Nvidia GPU — at the cost of being locked to one frozen model version per chip.

Read next: AI chip architectures guide · AMD Acquires Taalas · Anthropic hires Amir Salek (TPU founder) · GPT-5.6 Sol Ultrafast on Cerebras · OpenAI Apple rebuttal · AI token pricing explained


Performance claims are based on OpenAI's June 24, 2026 announcement. Final benchmark numbers, pricing, and availability details will be in the forthcoming technical report. This article is not financial advice.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Aug 19, 2026

Cerebras CS-4: The Wafer-Scale Chip Claiming 30x Faster AI Inference

Cerebras announced CS-4, the third generation of its wafer-scale AI accelerator, claiming up to 30x faster inference than GPU systems and a new modular "Nexus" rack architecture built to deploy at hyperscale. We break down what's actually new, how it stacks up against GPUs and rival inference chips like Taalas, and what the claims mean before independent benchmarks land.

Aug 7, 2026

AMD Acquires Taalas: The Chip That Etches Model Weights Into Silicon

AMD announced the acquisition of Taalas, a Toronto startup that etches LLM weights directly into silicon instead of storing them in HBM. Its test chip served Llama 3.1 8B at 16,960 tokens/second. We break down the architecture, the speed claims, and the real tradeoffs Hacker News flagged.

Aug 26, 2026

NVIDIA Dynamo Shadow Engine: 7s LLM Failover vs 5-Min Cold Start

NVIDIA's Aug 25, 2026 technical blog documents Shadow Engine Recovery in Dynamo: a hot standby inference worker on the same GPU sharing model weights through GPU Memory Service — failover in 7.3 seconds vs 283-second cold restarts in a deliberate fault test. explainx.ai explains when shadow engines beat checkpoint snapshots for self-hosted LLM SLOs.