explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the comparison at a glance
  • Who are the five?
  • What does it cost to run? A worked example
  • Benchmarks: what can and cannot be compared
  • Which should you use?
  • What to watch before you commit
  • Bottom line
  • Related reading on explainx.ai
← Back to blog

explainx / blog

Small Models Compared: Haiku 5.5 vs GPT-6 Luna vs DeepSeek vs Gemini Flash vs GLM

Model Comparison, Claude Haiku 5.5, GPT-6 Luna, DeepSeek, Small Models

Five cheap models compared on price, cache cost, openness and benchmarks: Claude Haiku 5.5, GPT-6 Luna, DeepSeek V4.1 Flash, Gemini 3.7 Flash, GLM-5.3-Flash.

Oct 8, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Small Models Compared: Haiku 5.5 vs GPT-6 Luna vs DeepSeek vs Gemini Flash vs GLM

Small models stopped being the "good enough" tier and became the place where most agent tokens get spent. On October 7, 2026, Anthropic released Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 output, the same list price OpenAI charges for GPT-6 Luna. Alongside them sit three strong alternatives: DeepSeek V4.1 Flash, Gemini 3.7 Flash and GLM-5.3-Flash.

This post compares all five on what can be compared honestly: list price, cache price, context, licensing, and the benchmarks that share a table. It then says which to use for subagents, coding, long context and self-hosting. Every number comes from our own earlier coverage or a vendor announcement, and we say when a figure is a claim.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the comparison at a glance

table · 6 cols
Claude Haiku 5.5GPT-6 LunaDeepSeek V4.1 FlashGemini 3.7 FlashGLM-5.3-Flash
Input per million$0.10 (up to 100k)$0.10about $0.15 to $0.22 off-peak$0.75$0.15
Output per million$0.50 (up to 100k)$0.50$0.66 off-peak$3.75$0.50
Cache read per million$0.01not in our sources$0.003 off-peaknot in our sources$0.03
Contextnot stated in our sourcesnot statedup to 1M1M1M
WeightsClosedClosedMIT openClosedMIT open
MultimodalComputer and browser useYes in ChatGPT productsNative multimodalMultimodalMultimodal input
Best forSubagents, browser use, summariesOpenAI-stack agents, volumeCheap cached agent loops, self-hostingGoogle stack, 1M contextOpen-weight agents, long context

Prices are per million tokens, list, for short prompts, from our earlier coverage. DeepSeek charges double at peak hours; Haiku 5.5 rises to $0.50 input and $2.50 output above 100k tokens; Gemini 3.7 Flash's $0.75 and $3.75 is an introductory rate through December 31, 2026 before rising to $1.50 and $7.50 on January 1, 2027. Our sources disagree on DeepSeek's off-peak cache-miss input ($0.22 in the beta coverage, $0.15 in later GA coverage), so check the live price page.

Who are the five?

Claude Haiku 5.5

Anthropic's new small model, released October 7, 2026. It is the first Haiku with adjustable effort levels, adds computer-use and browser-use support in the SDKs, and is pitched as a subagent next to Sonnet 5.5 and Opus 5.5. Anthropic says it costs about 75% less than Haiku 4.5 on average. Full details are in our Haiku 5.5 launch post.

GPT-6 Luna

OpenAI's cheap, fast tier, repriced on September 22, 2026 from $0.20 and $1.20 to $0.10 and $0.50. Batch output is $0.25 per million. Independent testing by Artificial Analysis found small regressions on two coding evaluations despite the price cut, which our GPT-6 Sol and Luna launch post covers. For why people like this tier, see Calvin French-Owen on small models.

DeepSeek V4.1 Flash

DeepSeek's MIT-licensed model, generally available on September 10, 2026, replacing the V4-Pro tier. It reports 1M context, native multimodality and an asymmetric design with 8B active parameters during prefill and 16B during decode. The headline for cost is cache: a 60% cut to cache-hit input, to $0.003 per million off-peak. See the V4.1 Flash beta coverage and the KV cache compression explainer.

Gemini 3.7 Flash

Google's Flash tier at $0.75 and $3.75 per million with a 1M context window and multimodal input, in our August comparison. It is the premium option here, priced about seven times Haiku and Luna on input. Its launch charts showed leads on AutomationBench and Code Arena against Grok 4.6 and Sonnet 5, and a deficit on DeepSWE. See Gemini 3.7 Flash against Grok 4.6, Sonnet 5 and GPT-5.6.

GLM-5.3-Flash

Z.ai's MIT-licensed model, a 320B-parameter, 18B-active mixture of experts with 1M context and multimodal input, at $0.15 and $0.50 with cached input at $0.03. Z.ai reports it ahead of Claude Opus 4.8 on GDPVal-AA v2 and DeepSWE. Details in our GLM-5.3-Flash launch post.

What does it cost to run? A worked example

Take a pipeline that sends 10 million input tokens and 2 million output tokens a day through short prompts, with no caching.

table · 4 cols
ModelInput costOutput costDaily total
Claude Haiku 5.5$1.00$1.00$2.00
GPT-6 Luna$1.00$1.00$2.00
GLM-5.3-Flash$1.50$1.00$2.50
DeepSeek V4.1 Flash (off-peak, $0.22 / $0.66)$2.20$1.32$3.52
Gemini 3.7 Flash (introductory rate)$7.50$7.50$15.00

Now add caching. If 9 of those 10 million input tokens are cache reads and the other million are fresh:

table · 5 cols
ModelCache readsFresh inputOutputDaily total
Claude Haiku 5.5$0.09$0.10$1.00$1.19
GLM-5.3-Flash$0.27$0.15$1.00$1.42
DeepSeek V4.1 Flash$0.03$0.22$1.32$1.57

We omit Luna and Gemini from the second table because we do not have their cache-read prices in our sources. The lesson holds anyway: once caching is on, output price decides the bill, not input. Haiku and GLM have the lowest output prices, and DeepSeek's near-zero cache read matters most when context is huge and reused often. For the mechanics, see our prompt caching guide, and for how the Sonnet-tier cache cut works, the Sonnet 5.5 cache-read cost math.

Benchmarks: what can and cannot be compared

Be careful here. These five were not evaluated in one place, so most cross-model numbers you will see are apples to oranges. Here is what is genuinely comparable.

Haiku 5.5 vs GPT-6 Luna (one table, vendor-run)

Anthropic's comparison table, which also includes Sonnet 5.5 and Haiku 4.5 for reference:

table · 4 cols
BenchmarkHaiku 5.5GPT-6 LunaSonnet 5.5 (reference)
GDPval-AA v2.1162014371840
AA-Briefcase v1.1157813361824
OSWorld 2.1 (offline subset)72.4%48.9%83.9%
Terminal-Bench 4.039.2%16.4%70.6%
FrontierCode 1.1 (Main)46.4%42.4%52.1% (xhigh)
Chartography, no tools46.4%29.1%61.6%

Haiku 5.5 leads on every row. It is Anthropic's own comparison, so treat it as a strong signal, not a verdict. The ceiling remains Sonnet: Haiku trails by more than 30 points on Terminal-Bench 4.0.

Independent evidence on Haiku 5.5

Vals AI reported Haiku 5.5 at 90.4% on Vibe Code Bench, third on that leaderboard, at about $6.07 per run. It also warned that the model used far more tokens than Haiku 4.5 and cost more per test on every benchmark both were run on. That is the single most important caution in this post: cheap per token is not cheap per task.

The others, kept separate

  • DeepSeek V4 Flash 0731, the earlier Flash tier, was verified by ARC Prize at 89.0% on ARC-AGI-1 for about $0.02 per task and 61.4% on ARC-AGI-2 for about $0.04. V4.1 Flash is the newer version and we do not have an equivalent verification, so do not read those figures as V4.1 results. See the ARC-AGI cost write-up.
  • GLM-5.3-Flash reports a GDPVal-AA v2 score of 1773 against Opus 4.8's 1582. That is version 2 of the benchmark; Haiku and Luna above are on v2.1, so the numbers do not sit on the same scale. These are Z.ai's figures.
  • Gemini 3.7 Flash reports AutomationBench 30.4% and Code Arena 1588 Elo in Google's launch charts, against Grok 4.6 and Sonnet 5, not against the models here.

If a single page tells you one of these five "wins," ask which benchmark version, which effort setting and who ran it.

Which should you use?

table · 3 cols
Your situationPickWhy
Subagents beside Sonnet or Opus on ClaudeHaiku 5.5Built for it, with effort levels and browser use
Subagents or bulk work on OpenAIGPT-6 LunaSame $0.10 and $0.50 price, easy to run in the OpenAI stack
Huge, repeatedly reused context on a budgetDeepSeek V4.1 Flash$0.003 cache-hit price and 1M context
You must self-host or control weightsDeepSeek V4.1 Flash or GLM-5.3-FlashBoth MIT licensed
Multimodal agent work with open weightsGLM-5.3-Flash1M context, multimodal input, MIT
Google ecosystem, 1M context, strong tool useGemini 3.7 FlashPay about 7x more for it; check that you need it
Browser or computer use at low costHaiku 5.5Anthropic reports OSWorld 72.4% for it
Hardest coding tasksNone of theseUse Sonnet 5.5, Opus 5.5 or a flagship; see our Sonnet vs Luna comparison

A practical routing pattern: use a flagship to plan, a small model for the many cheap steps, and escalate when the small model fails a check. Pair that with Haiku subagent effort levels and a cache-friendly prompt layout.

What to watch before you commit

  • Measure cost per completed task. Run the same 50 tasks on two models and total tokens, retries and failures. Sticker price is a starting point.
  • Watch the tokenizer. Haiku 5.5 uses an updated tokenizer that spends slightly more tokens per task than Haiku 4.5, and other vendors differ too.
  • Check the peak and the promo. DeepSeek charges more at peak hours, and Gemini's rate is introductory until December 31, 2026.
  • Re-check licensing. MIT weights are permissive, but serving a 320B-parameter model yourself is a real infrastructure cost; see our note on Qwen3.8-Flash-Next and big MoE models.
  • Mind data residency. Chinese-lab APIs and self-hosting have different compliance profiles than US cloud APIs.
  • Use an independent index. Artificial Analysis's Intelligence Index v4.2 is a better cross-model reference than vendor charts, though it does not fix harness differences.

Bottom line

At the cheap end, Haiku 5.5, Luna and GLM-5.3-Flash sit within a few cents of each other, DeepSeek V4.1 Flash is the cache-cost and open-weights play, and Gemini 3.7 Flash is the premium choice that must justify its price. On the one direct comparison we have, Haiku 5.5 beats Luna, but a vendor-run table and the token-use warning mean you should test before moving traffic. The model with the lowest bill is the one that finishes your task in the fewest tokens, and you only learn that by measuring.

Related reading on explainx.ai

  • Claude Haiku 5.5 launch: pricing, benchmarks and the Sonnet cache cut
  • GPT-6 Sol and Luna launch pricing
  • DeepSeek V4.1 Flash API beta and native multimodal
  • DeepSeek V4.1 Flash KV cache and HBM reduction
  • Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6
  • GLM-5.3-Flash: Ox Alpha unmasked
  • Sonnet 5.5 cache read cost math
  • Small models have arrived: Luna economics

Prices and benchmarks are as of October 8, 2026, from vendor announcements and our earlier coverage; they change often, and several are vendor claims. DeepSeek, Gemini and GLM figures come from earlier model versions where noted. Check each provider's pricing page before budgeting.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 3, 2026

Claude Opus 5.5 Tops the Epoch Capabilities Index With 167 — What the Score Means

Epoch AI's Capabilities Index puts Claude Opus 5.5 first among 253 tracked models at 167, a 0.84-point lead over GPT-6 Astra. The headline is real, but the margin is small and the benchmark mix matters more than the rank.

Oct 3, 2026

DwarfStar 4: Can You Run ds4 on Your Mac This Week?

Salvatore Sanfilippo's ds4 is back on the Hacker News front page. The useful question is whether your Mac, Spark, or Strix Halo can hold a project quant this week, and whether Claude Code or Codex can talk to it.

Oct 2, 2026

DeepSeek Harness Desktop: Should You Switch in 2026?

DeepSeek opened a worldwide public preview of DeepSeek Harness with an installable desktop app, the existing Web UI, and the same Cordis "everything is a plugin" runtime. This is the durable explainer: what DSH is, what the October 2026 desktop surface changes, and whether you should leave Pi, OpenCode, or Claude Code.