explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • The cost-economics argument: from $1 to $0.10 per run
  • The framework: IQ-180 work vs. token-spewer work
  • What people are debating: the Bitter Lesson pushback
  • What this means for your model-selection strategy
  • Related reading
← Back to blog

explainx / blog

Small Models Have Arrived: Calvin French-Owen on Luna Economics

Calvin French-Owen argues GPT-5.6 Luna and GLM 5.3 crossed a usefulness threshold, cutting his pet-eval cost from $1 to $0.10. Here is what changes.

Aug 28, 2026·11 min read·Yash Thakker
AI PricingModel SelectionGPT-5.6 LunaAgent EconomicsSmall Language Models
go deep
Small Models Have Arrived: Calvin French-Owen on Luna Economics

An essay from Calvin French-Owen — Segment co-founder, now on the Claude Code team at Anthropic — climbed to #2 on Hacker News on August 26, 2026, pulling 453 points and 203 comments. Its title states its thesis flatly: "Small Models Have Arrived." The argument isn't that small models beat frontier models. It's that they've crossed a usefulness threshold cheap enough to change what an AI product can afford to be.

That's a cost story, not just a benchmark story — and cost stories are where explainx.ai spends most of its time. This post walks through French-Owen's argument, the pet-eval numbers behind it, the pushback it drew on Hacker News, and what it should change about how you pick a model for a given task. For the mechanics of token pricing itself, start with explainx.ai's AI token pricing explained — this post assumes that groundwork and builds the "which model tier, when" decision on top of it.

TL;DR

table · 2 cols
QuestionAnswer
Who wrote it, and where?Calvin French-Owen (Segment co-founder, now Anthropic/Claude Code) on his personal blog, calv.info, August 26, 2026
How big was the reaction?#2 on Hacker News, 453 points, 203 comments
Which models is he talking about?OpenAI's GPT-5.6 Luna, and GLM 5.3 as another model he cites at the cost/capability Pareto frontier (via an Artificial Analysis chart)
What's the headline number?His personal research-agent "pet eval" cost about $1/run with the prior model generation, about $0.10/run with Luna
What's his framework for deciding when to use which tier?"IQ 180 work" (rare, novel, genius-tier) vs. "token spewer work" (fast, responsive, reliable, most of the job)
What's the consumer-AI implication?Token cost, not just growth mechanics, has been the missing ingredient blocking a wave of consumer AI companies
What's the pushback?A Bitter Lesson-based sub-thread argues scaled/general methods win long-term; counter-replies say that's not what the Bitter Lesson actually claims
What should builders do with this?Route by task type, not by habit — reserve frontier tiers for genuinely novel problems, default to Luna/GLM-5.3-tier models for the rest
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

The cost-economics argument: from $1 to $0.10 per run

French-Owen says he's spent weeks using GPT-5.6 Luna across his own codebase, email, and knowledge base, and describes it as "shockingly capable, fast, and smart." He reports regularly seeing throughput around 100 tokens per second, and says even fairly complicated research threads — the kind that search across thousands of emails — run up an API bill of only tens of cents.

He's careful to note he still reaches for the most expensive, most capable models — Fable 5, GPT-5.6 Sol — for his own coding work. That's the trap: if you only ever benchmark against frontier models on hard problems, it's easy to miss how far the cheap, fast tier has come on everything else.

His clearest evidence is a personal benchmark he calls his pet eval: prompt an agent to research a named person on the internet, figure out what news they'd want, build a personalized micro-site with today's top stories for them, and search Hacker News, Reddit, and Twitter/X along the way. It's a deliberately agentic, multi-tool, multi-search task — not a single completion.

With the previous generation of Sonnet-class models, that eval cost about $1 per run. French-Owen calls that untenable for a consumer product charging less than WSJ- or Economist-level subscription prices — you can't run a $1-per-request agent behind a $10/month plan and survive unit economics. With Luna, the same eval costs about $0.10 per run, with results he describes as decent. His reaction: "Now we're talking!"

That 10x swing is the whole essay in one data point. It's also exactly the kind of task-level cost audit explainx.ai has argued for in what running an AI agent actually costs per month — the sticker price per million tokens tells you little until you multiply it by tool calls, retries, and context replay across a real multi-turn task.

Why this matters for consumer AI specifically

French-Owen connects this to a question he says he's heard from investors: "It's weird we're not seeing more consumer AI companies. Why is that?" His answer is token cost, not lack of ideas. The classic consumer-app playbook — cheap website, viral growth, raise money, scale, monetize with ads later — assumes marginal cost per user is near zero. Once every request carries real inference cost, that assumption breaks, and the capital required to reach scale jumps accordingly.

Cheap-and-good-enough models don't eliminate that cost. They shrink it enough that the old playbook becomes viable again for a wider range of products. A 10x cost reduction on the compute line of a consumer AI product is the difference between "burn cash until ad revenue catches up" and "the unit economics work from day one."

GLM 5.3 gets cited in the essay as another model now sitting on that same cost/capability Pareto frontier, via a chart credited to artificialanalysis.ai — one more data point that this isn't a single-vendor story, it's a tier of the market maturing at once. explainx.ai has tracked this same frontier shift from the pricing side in OpenAI's GPT-5.6 Luna and Terra price cuts, where Luna's 80% July 2026 price cut put it at roughly one-sixth the cost per task of Claude Opus 5 (Low) on the Artificial Analysis Intelligence Index — the same chart family French-Owen is drawing on here.

The framework: IQ-180 work vs. token-spewer work

The most reusable idea in the essay didn't come from French-Owen originally — he attributes it to Peter Reinhardt, his ex-Segment co-founder, who has since raised over $100M for Charm Industrial and closed a Series A for Revoy. Reinhardt's framework splits most startup work into two buckets:

  • "IQ 180 work" — rare, genius-level novel problem solving. The kind of thinking that unblocks a company at a structural level.
  • "Token spewer work" — being ultra-responsive and pushing the ball forward across many fronts simultaneously: calls, nudging people, blocking and tackling, follow-through.

Reinhardt's own estimate: roughly 95% of his work is bucket two, even though he's clear his companies would fail without the bucket-one work also happening. It's not that IQ-180 work doesn't matter — it's that it's a small fraction of total volume, and most of the job is reliable, high-throughput execution rather than singular insight.

Map that directly onto model selection and you get a genuinely useful heuristic: reserve frontier-tier models — Fable 5, GPT-5.6 Sol, Opus-class reasoning — for the IQ-180 slice of your workload, and default to Luna/GLM-5.3-tier models for everything that's really token-spewer work: triage, drafting, routine research, status updates, first-pass classification, the volume tasks that make up most of what an agent actually does all day.

French-Owen's own prediction follows from this split. Demand for frontier-level models, he expects, keeps compounding for fields that genuinely need novel breakthroughs — engineering, hard science, model training itself. But demand for fast, cheap, good-enough models is "just about to take off," because most human-to-human interaction inside a company — with coworkers, vendors, customers — is the responsive, reliable-handling archetype, not genius-tier work. He points out that hiring itself already skews this way: most roles a company fills are for reliability and responsiveness, not for singular brilliance. What's still missing to make this real for business, in his view, is better harnesses, prompt-injection safety, and clearer roles and permissions for agents operating at that lower cost tier.

This is the same tension explainx.ai mapped from a benchmarking angle in the model-selection energy math nobody is doing: task-aware routing between model tiers can cut cost and energy substantially for a small quality trade-off — because most requests in a real workload were never the hard 5% to begin with.

What people are debating: the Bitter Lesson pushback

The Hacker News discussion split into a few substantive threads worth knowing about if you're deciding how much weight to put on this argument.

The Bitter Lesson objection. A large sub-thread invoked Rich Sutton's Bitter Lesson — the observation that general, scaled methods (search and learning) have historically beaten hand-tuned, specialized ones as compute grows — to argue that small or specialized models won't stay competitive with frontier models over time. If scale keeps winning, the thinking goes, today's "good enough" small model is just a temporary artifact of frontier models being briefly overpriced.

Counter-replies in the same thread pushed back on the premise, not just the conclusion. They pointed out the Bitter Lesson is specifically a claim about search-plus-learning beating hand-crafted heuristics — not a general rule that "always use the biggest model available." Cost, latency, and task-fit, the counter-argument goes, still meaningfully favor smaller models for narrow, well-defined tasks today, independent of whether frontier scaling continues. The concrete example cited: narrow specialist systems like the chess engine Stockfish still beat general-purpose LLMs at chess, despite LLMs being vastly larger and more broadly capable. Scale winning "in general" and scale winning "on this specific task, today" are different claims, and the thread didn't fully resolve which one the Bitter Lesson actually licenses.

Neither side landed a knockout in the discussion, and this piece isn't taking a side on it either — it's an open disagreement about how durable the "small models are good enough" advantage will be, not a settled fact either direction.

The practical counter-thread. Separately from the Bitter Lesson debate, several commenters reported already routing real work to small models successfully. One said Luna now handles roughly 90% of their code changes, with bigger models reserved only for genuinely complex problems — an informal version of the IQ-180/token-spewer split in practice. Another commenter described building their own daily personalized news site as a pet eval of their own, using it the same way French-Owen does — as a running, personal benchmark for whether small models have crossed the "good enough" line, rather than trusting a published leaderboard.

What this means for your model-selection strategy

None of this changes the mechanics of token pricing — output still costs more than input, caching still discounts repeated context, reasoning tokens are still often billed even when hidden from the response. If those terms aren't second nature yet, AI token pricing, explained without the pricing-page fog is the place to start.

What this essay changes is the decision layered on top of that pricing knowledge: which tier to route a given task to, by default. Three takeaways carry over directly into a model-selection strategy:

  1. Audit your workload by IQ-180 vs. token-spewer share before picking a default model. If most of your agent's calls are triage, retrieval, drafting, or routine classification — the token-spewer bucket — a Luna/GLM-5.3-tier default is very likely leaving money on the table by defaulting to frontier pricing for volume work.
  2. Reserve frontier tiers for the tasks that actually need them. Novel debugging, architecture decisions, and genuinely hard reasoning are the IQ-180 slice — that's where Fable 5, GPT-5.6 Sol, and Opus-class models earn their premium. explainx.ai's own comparison of Fable 5, Grok 4.6, GPT-5.6 Sol, and Qwen3.8-Max shows how sharply cost and quality diverge even inside the frontier tier itself.
  3. Cost per completed task, not cost per token, is the real unit. The $1-to-$0.10 pet-eval swing is a cost-per-task number driven by throughput and tool-call efficiency, not a sticker-price comparison. explainx.ai's own hands-on numbers on this exact model pairing are in Claude Sonnet 5 vs. GPT-5.6 Luna Max: cost comparison.

The consumer-AI framing is the highest-stakes part of the argument, and it's worth taking seriously even if you're not building a consumer product: if a 10x cost reduction on a single model generation changes whether a subscription business's unit economics work, then model selection isn't a technical footnote in a product plan — it's a line item that decides whether the business model survives contact with real usage.

Related reading

  • PhoneLLM: an open voice-agent model claiming GPT-5.6 Terra quality at 1/18th the cost — a sharper, task-specific version of this exact economics story: a small fine-tuned model undercutting a frontier API on cost and latency
  • AI token pricing, explained without the pricing-page fog
  • OpenAI cuts GPT-5.6 Luna 80%, Terra 20%
  • GPT-5.6 Sol now unifies ChatGPT chat; Luna goes unlimited for free users
  • Claude Sonnet 5 vs. GPT-5.6 Luna Max: cost comparison
  • Gemini 3.7 Flash vs. Grok 4.6 vs. Sonnet 5 vs. GPT-5.6: the real numbers
  • Fable 5 vs. Grok 4.6 vs. GPT-5.6 Sol vs. Qwen3.8-Max comparison
  • The model-selection energy math nobody is doing
  • What running an AI agent actually costs per month

Source: Calvin French-Owen, "Small Models Have Arrived," calv.info, August 26, 2026 — discussed on Hacker News (news.ycombinator.com), #2, 453 points, 203 comments as of publication.


This post summarizes and reacts to Calvin French-Owen's essay and the Hacker News discussion around it as of August 28, 2026. Pricing figures for GPT-5.6 Luna reflect OpenAI's public rate card as of that date; verify current pricing before budgeting production workloads. Opinions attributed to Calvin French-Owen, Peter Reinhardt, and HN commenters are theirs, not explainx.ai's editorial position.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 26, 2026

The Model-Selection Energy Math Nobody Is Doing

Research finds task-aware model selection can cut energy 27.8% for a 3.9% utility trade-off. Cheaper and greener are often the same inference optimization.

Aug 28, 2026

PhoneLLM: An Open Voice-Agent Model Claiming GPT-5.6 Terra Quality at 1/18th the Cost

Daily, the team behind the open-source Pipecat voice-AI framework, released PhoneLLM Alpha 1 on August 27, 2026 — an open-weights fine-tune of NVIDIA Nemotron 3 Nano claiming GPT-5.6 Terra-level quality on phone-support tasks at roughly 1/18th the cost and faster first-token latency. Here is what it actually claims, how the numbers stack up, and why this is an early alpha, not an independently verified benchmark.

Aug 26, 2026

ChatGPT Business Premium Seats: $100/Seat for Power Users

ChatGPT Business now has two seat tiers: Standard ($20–25/user) and Premium ($100–125/user) with 5× usage and no five-hour throttle. explainx.ai maps who actually needs Premium, how it compares to personal Pro/Max, and the per-seat math for small teams choosing between Business and API billing.