explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Why "outrageously small" is the right framing
  • What AMX actually is, and why it matters here
  • Built by handing Claude Code "a pile of tokens"
  • What's confirmed and what to check yourself
  • The three discoveries (and a fourth worth knowing)
  • 10K tok/s: design target vs measured throughput
  • Architecture: built from the machine, not the other way around
  • Four failures that had to be diagnosed first
  • When a CPU-only model beats reaching for a GPU
  • Built with Claude Code — what that actually meant
  • Honest limitations — read before you cite
  • What people are asking
  • Builder takeaways
  • Related on explainx.ai
← Back to blog

explainx / blog

Greg Diamos Built a 10K Tok/s CPU Model, and Found 3 Surprises

Small Language Models, CPU Inference, Claude Code, Edge AI, AI Efficiency

Greg Diamos used Claude Code to build a small CPU-only model for data processing at 10K tokens/sec, and published three findings about small neural nets worth revisiting. What amx-reasoning-v1 actually shows.

Sep 8, 2026·17 min read·Yash Thakker
add explainx.ai
go deep
Greg Diamos Built a 10K Tok/s CPU Model, and Found 3 Surprises

Not every notable AI story this week is about a frontier lab shipping a trillion-parameter model. Greg Diamos, who describes himself as building AI supercomputers, needed something much smaller and much faster: a CPU-only model that could process data at roughly 10,000 tokens per second with no GPU involved. His solution, built with the help of Claude Code, is amx-reasoning-v1-instruct — and he says the exercise surfaced three findings worth revisiting about "outrageously small neural nets."

TL;DR

table · 2 cols
QuestionAnswer
What did Diamos build?amx-reasoning-v1-instruct, a small CPU-only model for data processing
Target speedRoughly 10,000 tokens/sec on CPU
How was it built?Diamos gave Claude Code "a pile of tokens" to build it from scratch
Where's it published?Hugging Face — gdiamos/amx-reasoning-v1-instruct
What does "AMX" mean?Intel's Advanced Matrix Extensions — CPU instructions that accelerate matrix math
What's the headline claim?Three interesting discoveries about small neural nets, detailed in the paper
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Why "outrageously small" is the right framing

Diamos's framing — that the field should "revisit outrageously small neural nets" — cuts against the dominant 2026 narrative of ever-larger frontier models. But it's a genuinely useful counter-question: not every task needs a model with hundreds of billions of parameters and a GPU cluster behind it. A data-processing pipeline that needs to classify, extract, or transform text at high throughput often has a much narrower capability requirement than a general-purpose chat assistant — and a small model tuned specifically for that narrower task, running on CPU hardware that's already provisioned and idle, can be the more practical engineering choice.

This is the same logic behind explainx.ai's coverage of other small, purpose-built models this year — MiniCPM5-1B's tiny-model breakthrough, the ESP32 28.9M-parameter LLM running per-layer, and PrismML's Bonsai 27B 1-bit ternary model for phones — a consistent 2026 thread of teams asking "how small can this actually be for what I need," rather than defaulting to the largest available model.

What AMX actually is, and why it matters here

AMX — Advanced Matrix Extensions — is an Intel CPU instruction set specifically designed to accelerate matrix multiplication, the core mathematical operation underlying neural network inference. Naming the model after it is a signal about the engineering approach: rather than treating CPU inference as a fallback for when a GPU isn't available, amx-reasoning-v1 is reportedly built to actively exploit CPU-level matrix acceleration hardware most inference stacks don't specifically target. That's a meaningfully different design philosophy than running a GPU-optimized model on CPU as a compatibility afterthought — it's building for the hardware you actually have, rather than the hardware you wish you had.

Built by handing Claude Code "a pile of tokens"

The build process itself is worth noting: Diamos describes giving Claude Code a large amount of compute budget and letting it build the model, rather than hand-designing the architecture himself from scratch. That's consistent with the growing pattern explainx.ai has tracked of using coding agents not just to write application code, but to conduct genuine model-architecture experimentation — OpenAI's own "research intern" milestone documented a similar shift internally, where coding agents increasingly handle substantive research tasks rather than boilerplate.

What's confirmed and what to check yourself

Diamos's announcement thread teased "three interesting discoveries," but his full write-up and the Hugging Face model card go considerably deeper. Here is what the primary sources actually say — including one finding that did not fit in the tweet.

The three discoveries (and a fourth worth knowing)

1. Reasoning behaviours show up early

Diamos assumed in-context copying, positional manipulation, and simple arithmetic would need training budgets far beyond what one CPU core could afford. They did not.

At 259 million tokens — about nine hours on a single Intel Emerald Rapids core with OMP_NUM_THREADS=1 — a block-routed mixture-of-experts variant with 3.65M active parameters scored:

table · 3 cols
Probe taskScoreChance baseline
In-context induction93%7%
Positional shift90%7%
Two-digit addition47%23%

That induction result is thirteen times chance after roughly 5% of the tokens consumed in Diamos's longest run. The probes are synthetic and explicit — induct requires finding an earlier occurrence of the current token and emitting what followed; shift requires a positional offset with no content matching — so high scores imply an in-context circuit rather than memorised trivia.

Two measurement traps nearly invalidated the result. Early probes drew token IDs uniformly over the vocabulary; in an 18M-token corpus sample, a uniformly drawn ID has median frequency zero and never appears 83.5% of the time — measuring near-random embeddings, not circuits. Redrawing by corpus frequency moved measured OV rank from 42,138 to 152 on an identical checkpoint. An early eval also included four tasks not in the training mixture; all models scored 0% on those, making the overall score look broken when the eval was wrong.

2. The loss does not saturate

The longest pretraining run consumed 4.91 billion tokens on a 3.3M-active-parameter dense model — about four days on one core. That is 1,481 tokens per active parameter, roughly 74× the Chinchilla ratio.

Smoothed training loss fell monotonically within each phase and was still descending when the run ended for scheduling reasons, not convergence. A least-squares fit over the final 1B tokens shows slope −0.123 nats per billion tokens. Diamos argues this is deliberate: Chinchilla optimises for training compute; if you optimise for inference cost, you spend budget on tokens rather than parameters — affordable only when the model is small enough to run fast on CPU.

The fixed foundation validation set is noisy and roughly flat after ~370M tokens, so continued improvement shows most clearly in training loss and annealed validation sets — not every held-out curve they track.

3. Post-training compounds — the base model is unusable

The pretrained checkpoint is not weak at teacher forcing — 43.4% top-1, perplexity 19.8 — but it cannot generate. Free-running decode collapses into an absorbing repetition basin within about five tokens, with entropy 0.0005 and p(top-1) = 1.0000.

No decode-time trick fixes that distribution. Top-k, top-p, min-p, and temperature have nothing to reshape. A no-repeat 3-gram ban escapes the basin by construction but moves Dolly F1 from 0.3% to 1.0% — effectively zero to zero. Only training fixed it.

Four sequential post-training rounds each addressed a failure the previous round exposed:

table · 3 cols
RoundFailure addressedEffect
Instruction SFTCannot stop or generateStop-on-EOT 0% → 52.5%
QA SFTNo extractive skillEM 0 → 13.9%, but 57.7% over-abstention
Abstention rebalanceRefuses when it knowsOver-abstention 57.7% → 19.8%
Vocabulary fixJunk-token answersEM 15.3% → 18.2%, F1 19.6% → 23.2%

The shipped amx-reasoning-v1-instruct checkpoint is the product of that stack. On held-out extractive QA with greedy decoding and the vocabulary mask applied, it reaches 18.2% exact match and 23.2% F1 — weak by frontier standards, surprising for 3,315,552 active parameters trained end-to-end on one core. Applying the QA stage on top of instruction tuning rather than directly on the base was worth +1.3 EM, +2.2 F1, and 8.7 points less over-abstention.

4. Curated data acts like upstream distillation

Diamos's fourth finding — less viral than the first three, but arguably the most actionable — is that every natural-language and code source in the training mixture is a curated artifact built with large models. Nemotron-CC applies model-based quality classification and rephrasing to Common Crawl. Nemotron-CC-Math uses model-assisted extraction. OpenCodeReasoning is reasoning traces generated by large models and comprises three quarters of the anneal phase.

Training a 3.3M-active-parameter model on that mixture is distillation without a teacher at training time: large models already decided what text was worth keeping and wrote many of the traces upstream. Results establishing what tiny models "cannot do" in earlier eras may have measured data distribution, not hard capacity limits — because corpora like this did not exist when sub-10M models were last studied seriously.

The practical implication: a fixed one-core training budget may buy more capability every year as curation improves, with no architecture or hardware change. Diamos lists "hold the config fixed, re-run against newer corpus releases" as the cheapest experiment in the space.

10K tok/s: design target vs measured throughput

The tweet-friendly 10,000 tokens per second figure is a roofline-derived design target, not a published generation benchmark. Diamos is explicit about this in his write-up:

  • Measured end-to-end training throughput on one AMX core ranges from 6,616 tok/s (window MoE) to 14,572 tok/s (compute-matched dense) in the paper's Table 9 — forward plus backward, not inference-only micro-benchmarks.
  • Inference should be faster than training throughput, but measured generation latency has not been published yet.
  • Until it is, treat "10k tok/s on CPU" as arithmetic from the roofline, not a latency claim you can cite in a SLA.

The roofline itself is better than many builders assume. On one Intel Xeon Silver 4514Y core with AMX:

table · 2 cols
MetricValue
Best single bf16 GEMM2,231 GF/s (M=2048, K=256, N=1280)
bf16 at 1024³976–1,074 GF/s
fp32 at 1024³, no AMX121 GF/s
GEMM dispatch floor1.4–1.5 µs
L2 per core2 MiB

AMX delivers roughly 9× over AVX-512 fp32 at the shapes this model uses. Diamos's design rule: if active parameters are small enough that hot weights live in 2 MiB of L2, generation above 10,000 tok/s on one core becomes "just arithmetic" — provided you issue few large GEMMs rather than many small ones, because the dispatch floor makes call count a first-order cost.

That is why routing happens per block of 256 consecutive tokens rather than per token, and why selected experts get concatenated into a single matmul. A conventional per-token MoE issuing one small GEMM per token would spend its entire budget in dispatch overhead.

Architecture: built from the machine, not the other way around

amx-reasoning-v1-instruct is not a shrunken GPT clone dropped onto CPU as an afterthought. The architecture follows from the roofline:

table · 2 cols
FieldValue
d_model256
Layers6 (lin, swa, swa, lin, swa, lin)
MixersSliding-window attention (window 256), log-decay linear attention (d_state 32)
MLP width640
Vocabulary16,384
Stored parameters7,492,448
Active parameters per token3,315,552

Attention is confined to document boundaries within packed training windows — without isolation, sliding-window attention would leak across neighbours while linear attention carries state across the whole window.

You cannot load this checkpoint with AutoModelForCausalLM.from_pretrained — the architecture is not a standard transformers class. The Hugging Face repo ships the m2r source, example.py, and a generation.json vocabulary mask that is mandatory for correct decoding.

Four failures that had to be diagnosed first

None of the discoveries above was measurable until Diamos's team found four failures invisible in the loss curve:

  1. MoE collapse — 128 experts collapsed into one function; MoE layers contributed 0.00% of the residual stream while loss looked plausible. Diagnosis: participation ratio of RMSNorm gain read 1.0 on MoE layers vs 219–303 on dense layers at the same depth. A permutation ablation — randomly reassign experts, watch loss not move — would have caught this on day one.

  2. Zero-init expert insertion under top-k — inserting a new expert with zero keys is function-preserving in dense routing but not under top-k, because the new row still participates in ranking and can displace the k-th expert.

  3. Routing statistic dominance — switching the router input from prefix mean to windowed mean over the previous block alone moved flat validation from 4.431 to 4.041. Expert identity matters enormously; expert placement within a sequence costs 0.0000 nats when permuted — a mechanism they published alongside an honest note that their explanation is not yet fully tested.

  4. 147 vocabulary rows eating the decoder — tokens like ' ballo' and 'Frequently' (corpus count zero) kept logits near 0 while trained-but-wrong rows were pushed to ~−7.9, so the argmax fell through to untrained rows precisely when the model was uncertain. Before masking, those two tokens were 29% of all DROP answers. Masking 800 token IDs moved EM from 15.3% to 18.2%.

Diamos's builder advice from the thread: at this scale, print the generations. Scores said "weak model"; generations said "bug."

When a CPU-only model beats reaching for a GPU

Diamos's use case is deliberately boring: a huge dataset, no GPU cluster budget, simple extraction and classification at throughput high enough to process the whole corpus instead of sampling 1%.

Once you are at ~10k tok/s on CPU, matmul stops being the bottleneck and tokenization and I/O take over — which is also the point where running the model over the entire corpus becomes cheaper than deciding which slice to inspect. That is the same economic logic behind explainx.ai's edge and tiny-model coverage this year: MiniCPM5-1B pushing capability down, ESP32's 28.9M-parameter per-layer LLM for microcontrollers, and PrismML's phone-scale ternary models — all asking how small and how local inference can go for a defined task.

AMX is the unlock Diamos emphasises over raw core count. Intel held off shipping tensor-style units in mainstream CPUs for years; AMX on Emerald Rapids and later Xeons means hundreds of cores each carry a real matrix unit. This is not a return to the Knights Landing era without tensor cores — it is CPU inference built for the matrix hardware that now ships by default on server SKUs.

For Coral-style edge deployments and batch data pipelines, the comparison is not "can this beat GPT-5 on MMLU?" It is "can this run 500 cores overnight on hardware I already pay for and extract answers from text I would otherwise skip?"

Built with Claude Code — what that actually meant

The build process is part of the story. Diamos gave Claude Code a large token budget and directed it to construct the model and paper, checking in for roughly 15 minutes a day for a week. The paper lists Claude Opus 5 as a named co-author with its version string — not a generic "AI assisted" footnote — because the version makes the claim auditable.

The productive loop, by Diamos's account: Claude insisted the model was too small to learn behaviour X; Diamos asked whether X appeared in the dataset; isolation experiments ran; Claude responded "I can't believe it did." That disagreement loop produced results neither party would have reached alone — including Section 8.3's published refutation of their own stated mechanism when an ablation killed it.

That pattern mirrors explainx.ai's broader 2026 thread on coding agents doing substantive research work — OpenAI's research acceleration metrics documented a similar shift internally, where agent workdays increasingly cover architecture experimentation rather than boilerplate commits.

For reproducibility, Diamos's advice to builders is literal: paste the paper into Claude Code. Configurations and training logs ship with the Hugging Face repo; the paper is written to be rebuildable from them.

Honest limitations — read before you cite

The model card and paper are unusually frank about what does not work:

table · 2 cols
LimitationDetail
No arithmetic over passagesRetrieves and compares; fails subtraction across dates (e.g. "46 years" → outputs "4")
Weak abstention62.9% missed abstention on SQuAD v2 unanswerables; 16.1% over-abstention on answerable rows
Extractive onlyNeeds a supplied passage; multi-item list answers largely absent
No alignmentWill reproduce corpus biases; no safety or preference training
Small eval samplen=30 per synthetic task, one seed — supports learning specific operations, not broad generalisation
Equal-time MoE vs denseNot yet run under wall-clock budget; per-token MoE wins may not hold per-second

The add probe shares its answer space with evaluation by construction — Diamos reports the 47% addition number as the weakest of the three synthetic scores, not the headline claim.

Synthetic priming on repeated spans bought 4.1× on the synthetic task but essentially nothing on natural repetition in held-out documents. A circuit trained on random order fires on random order — transfer is not automatic.

What people are asking

Is 18.2% EM actually good? No — not against contemporary chat models. Yes — as evidence that 3.3M active parameters on one CPU core can do passage-grounded retrieval at all, with failures that are specific (arithmetic, list extraction, adversarial unanswerables) rather than diffuse nonsense.

Should I fine-tune my own 7M model on CPU? Only if your task matches the economics: high-volume, narrow extraction, hardware already provisioned, tolerance for multi-stage post-training. This is not a general assistant recipe.

How does this relate to fine-tuning guides? The four-round post-training stack is a case study in sequential SFT addressing distinct failure modes — stopping, extractive skill, abstention calibration, vocabulary hygiene — rather than one-shot instruction tuning.

Did Diamos use distillation loss? No teacher at training time and no distillation objective — but curated corpora built with large models upstream function as implicit distillation, which is the fourth finding.

Builder takeaways

  1. Measure the hardware first — derive architecture from roofline and dispatch floors, not from a GPU checkpoint shrunk blindly.
  2. Probe with care — eval design can fake broken models (wrong token sampling) or hide real circuits (tasks outside the training mixture).
  3. Plan post-training as a sequence — if round one of SFT disappoints, that diagnoses round one, not the base capacity.
  4. Mask or fix dead vocabulary rows — zero-count tokens in sampled softmax can dominate argmax when the model is uncertain.
  5. Print generations, not just loss — at small scale, qualitative output separates bugs from weakness faster than aggregate metrics.
  6. Separate training tok/s from inference latency — cite 6,616 tok/s as measured training throughput until generation benchmarks ship.
  7. Hold corpora constant, refresh curation — the cheapest experiment may be rerunning the same config on next year's Nemotron release.

The claim is not that amx-reasoning-v1-instruct is good. It is that this much — induction, shift, partial arithmetic, extractive QA — is reachable on one core, arrives much earlier in token budget than assumed, and that remaining failures are localised and nameable rather than proof that small models cannot reason at all.

Related on explainx.ai

  • FrogNano: Microsoft 4B SWE agent, 61.5% SWE-bench, no distillation (Sep 9, 2026) — another minimal-hardware agent bet, trained with RL on synthetic tasks only
  • MiniCPM5-1B: tiny AI model breakthrough
  • ESP32 AI: a 28.9M-parameter LLM running per-layer
  • PrismML Bonsai 27B: 1-bit ternary model for phones
  • Coral Edge AI platform: complete guide
  • OpenAI's research acceleration: 3.1 agent-workdays per human
  • What is fine-tuning an LLM? LoRA, QLoRA, SFT, RLHF explained

Sources

  • Greg Diamos on X, September 7, 2026
  • Outrageously Small Neural Networks — Greg Diamos blog, September 7, 2026
  • gdiamos/amx-reasoning-v1-instruct on Hugging Face
  • paper.pdf — primary evaluation source

This post reflects Greg Diamos's September 7, 2026 announcement, his follow-up blog write-up, and the Hugging Face model card as of September 11, 2026. Throughput figures distinguish measured training rates (6,616 tok/s) from the 10,000 tok/s roofline design target; generation latency benchmarks were not published at time of writing.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 30, 2026

Has Opus 5.5 Been Nerfed Yet? What Livenerf Can Say

Livenerf is an independent 30-day append-only benchmark for whether Opus 5.5 quietly degrades after launch. As of September 29 it is still collecting baseline. The first possible call is around October 24. This is what the instrument can and cannot see.

Sep 29, 2026

Claude Code Can Build Evals and Hillclimb Them: /claude-api build-eval

On September 28, 2026, Lance Martin published Anthropic’s playbook for eval design and hillclimbing in the claude-api skill. /claude-api build-eval interviews you, samples production traces, and validates graders. /claude-api hillclimb then patches one change per round against a train/test split. One internal support benchmark cut cost to about one-fifth while lifting held-out accuracy from 78.6% to 90.5%.

Sep 29, 2026

Claude Opus 5.5 vs Sonnet 5.5: Same Family, Different Bill

Anthropic’s launch table puts Sonnet 5.5 within a few points of Opus 5.5 on knowledge-work Elo and ahead on Terminal-Bench 4.0. Artificial Analysis still ranks Opus 58 vs Sonnet 56 — and Sonnet at max costs more per index task. This is the routing matrix for Claude Code and the API.