Small models stopped being the "good enough" tier and became the place where most agent tokens get spent. On October 7, 2026, Anthropic released Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 output, the same list price OpenAI charges for GPT-6 Luna. Alongside them sit three strong alternatives: DeepSeek V4.1 Flash, Gemini 3.7 Flash and GLM-5.3-Flash.
This post compares all five on what can be compared honestly: list price, cache price, context, licensing, and the benchmarks that share a table. It then says which to use for subagents, coding, long context and self-hosting. Every number comes from our own earlier coverage or a vendor announcement, and we say when a figure is a claim.
TL;DR: the comparison at a glance
| Claude Haiku 5.5 | GPT-6 Luna | DeepSeek V4.1 Flash | Gemini 3.7 Flash | GLM-5.3-Flash | |
|---|---|---|---|---|---|
| Input per million | $0.10 (up to 100k) | $0.10 | about $0.15 to $0.22 off-peak | $0.75 | $0.15 |
| Output per million | $0.50 (up to 100k) | $0.50 | $0.66 off-peak | $3.75 | $0.50 |
| Cache read per million | $0.01 | not in our sources | $0.003 off-peak | not in our sources | $0.03 |
| Context | not stated in our sources | not stated | up to 1M | 1M | 1M |
| Weights | Closed | Closed | MIT open | Closed | MIT open |
| Multimodal | Computer and browser use | Yes in ChatGPT products | Native multimodal | Multimodal | Multimodal input |
| Best for | Subagents, browser use, summaries | OpenAI-stack agents, volume | Cheap cached agent loops, self-hosting | Google stack, 1M context | Open-weight agents, long context |
Prices are per million tokens, list, for short prompts, from our earlier coverage. DeepSeek charges double at peak hours; Haiku 5.5 rises to $0.50 input and $2.50 output above 100k tokens; Gemini 3.7 Flash's $0.75 and $3.75 is an introductory rate through December 31, 2026 before rising to $1.50 and $7.50 on January 1, 2027. Our sources disagree on DeepSeek's off-peak cache-miss input ($0.22 in the beta coverage, $0.15 in later GA coverage), so check the live price page.
Who are the five?
Claude Haiku 5.5
Anthropic's new small model, released October 7, 2026. It is the first Haiku with adjustable effort levels, adds computer-use and browser-use support in the SDKs, and is pitched as a subagent next to Sonnet 5.5 and Opus 5.5. Anthropic says it costs about 75% less than Haiku 4.5 on average. Full details are in our Haiku 5.5 launch post.
GPT-6 Luna
OpenAI's cheap, fast tier, repriced on September 22, 2026 from $0.20 and $1.20 to $0.10 and $0.50. Batch output is $0.25 per million. Independent testing by Artificial Analysis found small regressions on two coding evaluations despite the price cut, which our GPT-6 Sol and Luna launch post covers. For why people like this tier, see Calvin French-Owen on small models.
DeepSeek V4.1 Flash
DeepSeek's MIT-licensed model, generally available on September 10, 2026, replacing the V4-Pro tier. It reports 1M context, native multimodality and an asymmetric design with 8B active parameters during prefill and 16B during decode. The headline for cost is cache: a 60% cut to cache-hit input, to $0.003 per million off-peak. See the V4.1 Flash beta coverage and the KV cache compression explainer.
Gemini 3.7 Flash
Google's Flash tier at $0.75 and $3.75 per million with a 1M context window and multimodal input, in our August comparison. It is the premium option here, priced about seven times Haiku and Luna on input. Its launch charts showed leads on AutomationBench and Code Arena against Grok 4.6 and Sonnet 5, and a deficit on DeepSWE. See Gemini 3.7 Flash against Grok 4.6, Sonnet 5 and GPT-5.6.
GLM-5.3-Flash
Z.ai's MIT-licensed model, a 320B-parameter, 18B-active mixture of experts with 1M context and multimodal input, at $0.15 and $0.50 with cached input at $0.03. Z.ai reports it ahead of Claude Opus 4.8 on GDPVal-AA v2 and DeepSWE. Details in our GLM-5.3-Flash launch post.
What does it cost to run? A worked example
Take a pipeline that sends 10 million input tokens and 2 million output tokens a day through short prompts, with no caching.
| Model | Input cost | Output cost | Daily total |
|---|---|---|---|
| Claude Haiku 5.5 | $1.00 | $1.00 | $2.00 |
| GPT-6 Luna | $1.00 | $1.00 | $2.00 |
| GLM-5.3-Flash | $1.50 | $1.00 | $2.50 |
| DeepSeek V4.1 Flash (off-peak, $0.22 / $0.66) | $2.20 | $1.32 | $3.52 |
| Gemini 3.7 Flash (introductory rate) | $7.50 | $7.50 | $15.00 |
Now add caching. If 9 of those 10 million input tokens are cache reads and the other million are fresh:
| Model | Cache reads | Fresh input | Output | Daily total |
|---|---|---|---|---|
| Claude Haiku 5.5 | $0.09 | $0.10 | $1.00 | $1.19 |
| GLM-5.3-Flash | $0.27 | $0.15 | $1.00 | $1.42 |
| DeepSeek V4.1 Flash | $0.03 | $0.22 | $1.32 | $1.57 |
We omit Luna and Gemini from the second table because we do not have their cache-read prices in our sources. The lesson holds anyway: once caching is on, output price decides the bill, not input. Haiku and GLM have the lowest output prices, and DeepSeek's near-zero cache read matters most when context is huge and reused often. For the mechanics, see our prompt caching guide, and for how the Sonnet-tier cache cut works, the Sonnet 5.5 cache-read cost math.
Benchmarks: what can and cannot be compared
Be careful here. These five were not evaluated in one place, so most cross-model numbers you will see are apples to oranges. Here is what is genuinely comparable.
Haiku 5.5 vs GPT-6 Luna (one table, vendor-run)
Anthropic's comparison table, which also includes Sonnet 5.5 and Haiku 4.5 for reference:
| Benchmark | Haiku 5.5 | GPT-6 Luna | Sonnet 5.5 (reference) |
|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 1336 | 1824 |
| OSWorld 2.1 (offline subset) | 72.4% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | 42.4% | 52.1% (xhigh) |
| Chartography, no tools | 46.4% | 29.1% | 61.6% |
Haiku 5.5 leads on every row. It is Anthropic's own comparison, so treat it as a strong signal, not a verdict. The ceiling remains Sonnet: Haiku trails by more than 30 points on Terminal-Bench 4.0.
Independent evidence on Haiku 5.5
Vals AI reported Haiku 5.5 at 90.4% on Vibe Code Bench, third on that leaderboard, at about $6.07 per run. It also warned that the model used far more tokens than Haiku 4.5 and cost more per test on every benchmark both were run on. That is the single most important caution in this post: cheap per token is not cheap per task.
The others, kept separate
- DeepSeek V4 Flash 0731, the earlier Flash tier, was verified by ARC Prize at 89.0% on ARC-AGI-1 for about $0.02 per task and 61.4% on ARC-AGI-2 for about $0.04. V4.1 Flash is the newer version and we do not have an equivalent verification, so do not read those figures as V4.1 results. See the ARC-AGI cost write-up.
- GLM-5.3-Flash reports a GDPVal-AA v2 score of 1773 against Opus 4.8's 1582. That is version 2 of the benchmark; Haiku and Luna above are on v2.1, so the numbers do not sit on the same scale. These are Z.ai's figures.
- Gemini 3.7 Flash reports AutomationBench 30.4% and Code Arena 1588 Elo in Google's launch charts, against Grok 4.6 and Sonnet 5, not against the models here.
If a single page tells you one of these five "wins," ask which benchmark version, which effort setting and who ran it.
Which should you use?
| Your situation | Pick | Why |
|---|---|---|
| Subagents beside Sonnet or Opus on Claude | Haiku 5.5 | Built for it, with effort levels and browser use |
| Subagents or bulk work on OpenAI | GPT-6 Luna | Same $0.10 and $0.50 price, easy to run in the OpenAI stack |
| Huge, repeatedly reused context on a budget | DeepSeek V4.1 Flash | $0.003 cache-hit price and 1M context |
| You must self-host or control weights | DeepSeek V4.1 Flash or GLM-5.3-Flash | Both MIT licensed |
| Multimodal agent work with open weights | GLM-5.3-Flash | 1M context, multimodal input, MIT |
| Google ecosystem, 1M context, strong tool use | Gemini 3.7 Flash | Pay about 7x more for it; check that you need it |
| Browser or computer use at low cost | Haiku 5.5 | Anthropic reports OSWorld 72.4% for it |
| Hardest coding tasks | None of these | Use Sonnet 5.5, Opus 5.5 or a flagship; see our Sonnet vs Luna comparison |
A practical routing pattern: use a flagship to plan, a small model for the many cheap steps, and escalate when the small model fails a check. Pair that with Haiku subagent effort levels and a cache-friendly prompt layout.
What to watch before you commit
- Measure cost per completed task. Run the same 50 tasks on two models and total tokens, retries and failures. Sticker price is a starting point.
- Watch the tokenizer. Haiku 5.5 uses an updated tokenizer that spends slightly more tokens per task than Haiku 4.5, and other vendors differ too.
- Check the peak and the promo. DeepSeek charges more at peak hours, and Gemini's rate is introductory until December 31, 2026.
- Re-check licensing. MIT weights are permissive, but serving a 320B-parameter model yourself is a real infrastructure cost; see our note on Qwen3.8-Flash-Next and big MoE models.
- Mind data residency. Chinese-lab APIs and self-hosting have different compliance profiles than US cloud APIs.
- Use an independent index. Artificial Analysis's Intelligence Index v4.2 is a better cross-model reference than vendor charts, though it does not fix harness differences.
Bottom line
At the cheap end, Haiku 5.5, Luna and GLM-5.3-Flash sit within a few cents of each other, DeepSeek V4.1 Flash is the cache-cost and open-weights play, and Gemini 3.7 Flash is the premium choice that must justify its price. On the one direct comparison we have, Haiku 5.5 beats Luna, but a vendor-run table and the token-use warning mean you should test before moving traffic. The model with the lowest bill is the one that finishes your task in the fewest tokens, and you only learn that by measuring.
Related reading on explainx.ai
- Claude Haiku 5.5 launch: pricing, benchmarks and the Sonnet cache cut
- GPT-6 Sol and Luna launch pricing
- DeepSeek V4.1 Flash API beta and native multimodal
- DeepSeek V4.1 Flash KV cache and HBM reduction
- Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6
- GLM-5.3-Flash: Ox Alpha unmasked
- Sonnet 5.5 cache read cost math
- Small models have arrived: Luna economics
Prices and benchmarks are as of October 8, 2026, from vendor announcements and our earlier coverage; they change often, and several are vendor claims. DeepSeek, Gemini and GLM figures come from earlier model versions where noted. Check each provider's pricing page before budgeting.
