A GitHub project called Strata hit the front page of Hacker News on October 5, 2026 with a claim that sounds like a typo: run Qwen3.8-Flash-Next, a 125-billion-parameter model, on a normal gaming PC with a 12 GB graphics card, at speeds between roughly 50 and 120 tokens per second. The thread collected 617 points and 289 comments, and the repository stands at about 11,200 stars and 980 forks at the time of writing.
This post covers what Strata actually does, the speed numbers and what you need to reproduce them, which quantized size to pick, and where the community pushed back. The short version: the speed is real on the hardware people reported, the quality story depends heavily on your task, and nobody should treat a 2-bit result as equal to a hosted frontier model.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | An MIT-licensed engine plus one-click installer that runs Qwen3.8-Flash-Next (125B MoE) on one consumer GPU. |
| What hardware? | NVIDIA RTX 20-50 or supported AMD card with 12 GB+ VRAM, 32 GB+ RAM, ~80 GB disk. |
| How fast? | 53-94 tok/s on an RTX 5070; 44-60 tok/s on an RX 9070 XT; ~124 tok/s reported on an RTX 4090 plus 128 GB RAM. |
| Is it free? | Yes. MIT license; the model has its own license. |
| What quality loss? | Real. It uses roughly 2-3 bit quants. Coding holds up for many users, vision and long-tail tasks show degradation. |
| Does it work with Claude Code or Codex? | It serves OpenAI, Anthropic and Responses API endpoints on localhost, so yes, as a local backend. |
| Should I replace my cloud model? | For private, well-scoped coding, possibly. For hard long-horizon work, test first. |
What Strata is, and why it can fit 125B on 12 GB
The model is a mixture-of-experts: 24,576 small "experts," of which each token needs only 10. HN commenters put the active parameter count at roughly 6B per token. Strata exploits that sparsity with a three-tier memory plan, described in the repository's README as a kitchen:
- GPU (12-24 GB): keeps the few thousand most-used experts, the "counter."
- System RAM (32-64 GB): holds all experts, and the CPU works on the rest in parallel, the "pantry."
- SSD: holds a lookup table, and for the largest quants, streams part of the model during decoding.
On top of that sits speculative decoding: a small helper model guesses the next few tokens and the big model verifies them in one pass, which the README says gives the same answer 1.6-1.8x sooner. Long prompts are read in chunks of up to 8,192 tokens at over 1,000 tokens per second.
If you have read our explainer on what llama.cpp is, the idea will be familiar: GGUF-style quantized weights plus smart offload. Strata is a specialized engine for one architecture rather than a general runtime, which is both why it is fast and why several commenters asked why the work is not upstream. The main objection, voiced in the thread, is that upstream projects have strict contribution rules about human-understood code, so heavily AI-assisted engines end up as separate projects.
The speed table (and what it was measured on)
These are the README's own numbers for 4K-token answers and 32K-token prompts. The NVIDIA box is an RTX 5070 (12 GB) with a Ryzen 5 7600 and 64 GB RAM. The AMD box is an RX 9070 XT (16 GB) with a Ryzen 9 3900X and 47 GB RAM.
| Size | RTX 5070: writes | RTX 5070: reads prompt | RX 9070 XT: writes | RX 9070 XT: reads prompt |
|---|---|---|---|---|
| Q2_0 | 94 tok/s | 2,650 tok/s | 60 tok/s | 1,160 tok/s |
| IQ2_XS | 79 tok/s | 2,090 tok/s | 52 tok/s | 1,110 tok/s |
| IQ3_XXS | 62 tok/s | 1,750 tok/s | n/a | n/a |
| IQ3_S | 53 tok/s | 1,620 tok/s | n/a | n/a |
| Coder | 55 tok/s | 2,180 tok/s | 44 tok/s | 1,420 tok/s |
Community reports in the Hacker News thread, all self-reported and unverified:
- RTX 4090 + 128 GB DDR5: about 124 tokens per second, and another user above 110 with 3-token multi-token prediction on a 4090 with 128 GB DDR4.
- RTX 3090: about 60 tokens per second on IQ3_S; another user, on a power-capped 3090 with 64 GB DDR4, tuned to 40-60.
- RTX 3080 (10 GB) + 48 GB RAM, Coder variant: around 30 tokens per second, with other programs running.
- Radeon RX 9070 XT, IQ2: about 65 tokens per second.
- RTX 2060 (8 GB), Q2_0 at 128K context: 33 tokens per second decode and about 600 tokens per second prompt processing.
- No GPU at all (a laptop CPU with 96 GB of RAM, using plain llama.cpp rather than Strata): about 7 tokens per second.
The README itself says a 24 GB card such as the RTX 3090 "should write about 100-140 tokens per second." One practical note: by default Strata serves one request at a time. A parallel setting exists, but on a 12 GB card it makes each answer slower.
Which size should you pick?
The installer recommends a size by RAM. The README's table, in our words:
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | It fits, and it is built for code |
| 48 GB | IQ2_XS (or Q2_0 for speed) | Larger sizes do not fit |
| 64 GB | IQ2_XS, IQ3_XXS or IQ3_S | Everything fits; IQ3_S is best and slowest |
| 96 GB+ | IQ3_S or Unsloth UD-IQ4_XS | Room for the largest sizes |
Two caveats matter. First, Coder is a pruned model: half the experts removed. The README says it reaches 91% of the full model's SWE-bench Verified score, as measured by its authors, but it is weaker outside code, including Chinese and other CJK text. A Hacker News commenter who tried it for long coding sessions advised using the normal sizes instead. Second, the ~4-bit UD-IQ4_XS is a 94 GB download, and with less than about 80 GB of RAM Strata reads part of it from the SSD while answering, so it is slower on mid-range machines.
The quality debate: what the thread actually showed
This is where the thread got useful. Three camps emerged.
Skeptics of sub-4-bit. Several developers argued that 4-bit is the floor for reliable use, that headline tokens-per-second numbers rarely come with accuracy data, and that "2-bit and half the experts removed" is a different product from the full model. One commenter cited a recent paper finding 4-bit usually preserves performance while 2-bit often causes broad degradation. That is a fair prior, not a verdict on this specific, selectively quantized model.
Users with positive results. Several reported IQ3_XXS or IQ3_S performing well. One commenter said they were halfway through the DeepSWE coding benchmark on a 3-bit quant and found it roughly on par with older Claude Opus and Sonnet versions. Another ran a private suite and scored Strata at 89.1% against 71.9% for a competing 3090-optimized Qwen3.8-27B setup, though with a longer runtime. These are single-user results, not controlled studies, and we have not reproduced them.
One hard counterexample on vision. A user ran a 50-image coordinate-pointing test with the same GGUF and vision adapter on Strata and on llama.cpp. Strata's median error was 154.8 pixels versus 46.5 on llama.cpp. If you plan to use image input, test it against a reference before trusting it. The README itself notes that AMD cards read pictures through the CPU on Linux and cannot yet on Windows.
A few other honest limitations from the thread: the first message of a long chat is read in full (about one minute per 30,000 tokens), loading the model can make a PC unresponsive for one to three minutes, and thinking on this model can be long.
Is it better than Qwen3.8-27B?
Opinions split sharply, which is itself informative. Some developers said Flash-Next is clearly better for coding and finding bugs in generated code. Another said they found the 27B more accurate at FP8 and saw spelling errors in Flash-Next near 150K context, likely a configuration issue. The practical logic from the thread: a dense 27B at 4-bit is about 16 GB of weights read per token, while Flash-Next reads a much smaller active slice, so MoE can win on speed even at heavier compression, but only testing on your own tasks settles quality. Our earlier Qwen3.8-27B coverage and the Flash-Next release post give the model-level context.
How to try it today
- Check your headroom. 12 GB VRAM, at least 32 GB RAM, about 80 GB free disk, current NVIDIA or AMD driver.
- Get the repo. Clone or unzip Strata. On Windows run
START-HERE.bat; on Linux run./setup.sh. - Accept the recommended answers. The installer picks the model size by RAM and downloads about 70 GB. It resumes if interrupted.
- Wait out the first load. Expect a slow or frozen-looking PC for one to three minutes while it loads 35-55 GB.
- Connect your tools. Add an "OpenAI-compatible" provider at
http://127.0.0.1:8080/v1. For Anthropic-style clients, use/v1/messages(Claude Code readsANTHROPIC_BASE_URL=http://127.0.0.1:8080). Codex CLI uses/v1/responses. - Set thinking low for speed, high for hard problems.
If you serve it beyond your own machine, the README says always set an API key. For wiring local models into a coding agent, see our guides to running open models locally with OpenCode and fixing local models that loop.
The "let your AI set it up" prompt
Strata's README suggests pasting a one-line prompt into a coding assistant such as Claude Code, Cursor or Codex: set up Strata for me and follow docs/AI_SETUP.md. Hacker News predictably argued about it, comparing it to piping a script into bash. The honest answer is that it is the same trust decision as any installer: read what the agent will run. Commenters did exactly that, reading setup.py and setup.sh before trusting them. The repository also lists the Claude bot among its contributors, and one commenter flagged it, which fits the heavily AI-assisted engineering pattern others noted for inference engines generally.
What this means for what you build or pay
If your cloud coding bill is dominated by well-scoped tasks (refactors, tests, scripted edits) and you already own a 12-24 GB GPU and 64 GB of RAM, Strata is a free experiment worth an afternoon. The upside is privacy, no quota resets to wait for, and a flat hardware cost. The downside is quality you must measure rather than assume. A sensible pattern is routing easy, private jobs to the local model and keeping a hosted model for long-horizon work, the same two-harness hedge we recommend in our Codex 28-day pledge post. For hardware choices, compare against our MacBook vs dedicated GPU guide before buying.
The bigger signal is architectural. Models built on the Qwen4-style sparse design, which Alibaba previewed in Qwen3.8-Flash, reward engines that know where each expert lives. If that holds, expect more model-specific engines and faster consumer-grade local inference.
Related reading
- Qwen3.8-Flash-Next: the 125B MoE Alibaba teased
- Qwen3.8-Flash on QwenCloud and the Qwen4 preview
- OpenCode adds Qwen3.8-Flash 125B
- Qwen3.8-27B: the local model Hacker News put at #1
- What is llama.cpp?
- How to run open models locally with OpenCode
- MacBook vs dedicated GPU for local LLMs
- PrismML Bonsai 27B: 1-bit Qwen on a phone
Primary: the Strata GitHub README (Niko1221/Strata, v0.1.39) · the October 5, 2026 Hacker News discussion
Figures are as of October 5, 2026 and come from the project's README and self-reported community posts. Speeds depend on RAM, PCIe lanes, quant and context length; we have not independently benchmarked Strata. Quantized model quality varies by task, so test on your own workload.
