explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What Strata is, and why it can fit 125B on 12 GB
  • The speed table (and what it was measured on)
  • Which size should you pick?
  • The quality debate: what the thread actually showed
  • Is it better than Qwen3.8-27B?
  • How to try it today
  • The "let your AI set it up" prompt
  • What this means for what you build or pay
  • Related reading
← Back to blog

explainx / blog

Strata Runs Qwen3.8-Flash-Next 125B on a 12 GB Gaming GPU: Speed vs Quality

Local LLMs, Qwen, Open Source, Inference, Guides

Strata runs the 125B Qwen3.8-Flash-Next on one 12-24 GB GPU plus 64 GB RAM. Real speeds, which quant to pick, and what the HN thread says about quality.

Oct 5, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Strata Runs Qwen3.8-Flash-Next 125B on a 12 GB Gaming GPU: Speed vs Quality

A GitHub project called Strata hit the front page of Hacker News on October 5, 2026 with a claim that sounds like a typo: run Qwen3.8-Flash-Next, a 125-billion-parameter model, on a normal gaming PC with a 12 GB graphics card, at speeds between roughly 50 and 120 tokens per second. The thread collected 617 points and 289 comments, and the repository stands at about 11,200 stars and 980 forks at the time of writing.

This post covers what Strata actually does, the speed numbers and what you need to reproduce them, which quantized size to pick, and where the community pushed back. The short version: the speed is real on the hardware people reported, the quality story depends heavily on your task, and nobody should treat a 2-bit result as equal to a hosted frontier model.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?An MIT-licensed engine plus one-click installer that runs Qwen3.8-Flash-Next (125B MoE) on one consumer GPU.
What hardware?NVIDIA RTX 20-50 or supported AMD card with 12 GB+ VRAM, 32 GB+ RAM, ~80 GB disk.
How fast?53-94 tok/s on an RTX 5070; 44-60 tok/s on an RX 9070 XT; ~124 tok/s reported on an RTX 4090 plus 128 GB RAM.
Is it free?Yes. MIT license; the model has its own license.
What quality loss?Real. It uses roughly 2-3 bit quants. Coding holds up for many users, vision and long-tail tasks show degradation.
Does it work with Claude Code or Codex?It serves OpenAI, Anthropic and Responses API endpoints on localhost, so yes, as a local backend.
Should I replace my cloud model?For private, well-scoped coding, possibly. For hard long-horizon work, test first.

What Strata is, and why it can fit 125B on 12 GB

The model is a mixture-of-experts: 24,576 small "experts," of which each token needs only 10. HN commenters put the active parameter count at roughly 6B per token. Strata exploits that sparsity with a three-tier memory plan, described in the repository's README as a kitchen:

  • GPU (12-24 GB): keeps the few thousand most-used experts, the "counter."
  • System RAM (32-64 GB): holds all experts, and the CPU works on the rest in parallel, the "pantry."
  • SSD: holds a lookup table, and for the largest quants, streams part of the model during decoding.

On top of that sits speculative decoding: a small helper model guesses the next few tokens and the big model verifies them in one pass, which the README says gives the same answer 1.6-1.8x sooner. Long prompts are read in chunks of up to 8,192 tokens at over 1,000 tokens per second.

If you have read our explainer on what llama.cpp is, the idea will be familiar: GGUF-style quantized weights plus smart offload. Strata is a specialized engine for one architecture rather than a general runtime, which is both why it is fast and why several commenters asked why the work is not upstream. The main objection, voiced in the thread, is that upstream projects have strict contribution rules about human-understood code, so heavily AI-assisted engines end up as separate projects.

The speed table (and what it was measured on)

These are the README's own numbers for 4K-token answers and 32K-token prompts. The NVIDIA box is an RTX 5070 (12 GB) with a Ryzen 5 7600 and 64 GB RAM. The AMD box is an RX 9070 XT (16 GB) with a Ryzen 9 3900X and 47 GB RAM.

table · 5 cols
SizeRTX 5070: writesRTX 5070: reads promptRX 9070 XT: writesRX 9070 XT: reads prompt
Q2_094 tok/s2,650 tok/s60 tok/s1,160 tok/s
IQ2_XS79 tok/s2,090 tok/s52 tok/s1,110 tok/s
IQ3_XXS62 tok/s1,750 tok/sn/an/a
IQ3_S53 tok/s1,620 tok/sn/an/a
Coder55 tok/s2,180 tok/s44 tok/s1,420 tok/s

Community reports in the Hacker News thread, all self-reported and unverified:

  • RTX 4090 + 128 GB DDR5: about 124 tokens per second, and another user above 110 with 3-token multi-token prediction on a 4090 with 128 GB DDR4.
  • RTX 3090: about 60 tokens per second on IQ3_S; another user, on a power-capped 3090 with 64 GB DDR4, tuned to 40-60.
  • RTX 3080 (10 GB) + 48 GB RAM, Coder variant: around 30 tokens per second, with other programs running.
  • Radeon RX 9070 XT, IQ2: about 65 tokens per second.
  • RTX 2060 (8 GB), Q2_0 at 128K context: 33 tokens per second decode and about 600 tokens per second prompt processing.
  • No GPU at all (a laptop CPU with 96 GB of RAM, using plain llama.cpp rather than Strata): about 7 tokens per second.

The README itself says a 24 GB card such as the RTX 3090 "should write about 100-140 tokens per second." One practical note: by default Strata serves one request at a time. A parallel setting exists, but on a 12 GB card it makes each answer slower.

Which size should you pick?

The installer recommends a size by RAM. The README's table, in our words:

table · 3 cols
Your RAMTakeWhy
32 GBCoderIt fits, and it is built for code
48 GBIQ2_XS (or Q2_0 for speed)Larger sizes do not fit
64 GBIQ2_XS, IQ3_XXS or IQ3_SEverything fits; IQ3_S is best and slowest
96 GB+IQ3_S or Unsloth UD-IQ4_XSRoom for the largest sizes

Two caveats matter. First, Coder is a pruned model: half the experts removed. The README says it reaches 91% of the full model's SWE-bench Verified score, as measured by its authors, but it is weaker outside code, including Chinese and other CJK text. A Hacker News commenter who tried it for long coding sessions advised using the normal sizes instead. Second, the ~4-bit UD-IQ4_XS is a 94 GB download, and with less than about 80 GB of RAM Strata reads part of it from the SSD while answering, so it is slower on mid-range machines.

The quality debate: what the thread actually showed

This is where the thread got useful. Three camps emerged.

Skeptics of sub-4-bit. Several developers argued that 4-bit is the floor for reliable use, that headline tokens-per-second numbers rarely come with accuracy data, and that "2-bit and half the experts removed" is a different product from the full model. One commenter cited a recent paper finding 4-bit usually preserves performance while 2-bit often causes broad degradation. That is a fair prior, not a verdict on this specific, selectively quantized model.

Users with positive results. Several reported IQ3_XXS or IQ3_S performing well. One commenter said they were halfway through the DeepSWE coding benchmark on a 3-bit quant and found it roughly on par with older Claude Opus and Sonnet versions. Another ran a private suite and scored Strata at 89.1% against 71.9% for a competing 3090-optimized Qwen3.8-27B setup, though with a longer runtime. These are single-user results, not controlled studies, and we have not reproduced them.

One hard counterexample on vision. A user ran a 50-image coordinate-pointing test with the same GGUF and vision adapter on Strata and on llama.cpp. Strata's median error was 154.8 pixels versus 46.5 on llama.cpp. If you plan to use image input, test it against a reference before trusting it. The README itself notes that AMD cards read pictures through the CPU on Linux and cannot yet on Windows.

A few other honest limitations from the thread: the first message of a long chat is read in full (about one minute per 30,000 tokens), loading the model can make a PC unresponsive for one to three minutes, and thinking on this model can be long.

Is it better than Qwen3.8-27B?

Opinions split sharply, which is itself informative. Some developers said Flash-Next is clearly better for coding and finding bugs in generated code. Another said they found the 27B more accurate at FP8 and saw spelling errors in Flash-Next near 150K context, likely a configuration issue. The practical logic from the thread: a dense 27B at 4-bit is about 16 GB of weights read per token, while Flash-Next reads a much smaller active slice, so MoE can win on speed even at heavier compression, but only testing on your own tasks settles quality. Our earlier Qwen3.8-27B coverage and the Flash-Next release post give the model-level context.

How to try it today

  1. Check your headroom. 12 GB VRAM, at least 32 GB RAM, about 80 GB free disk, current NVIDIA or AMD driver.
  2. Get the repo. Clone or unzip Strata. On Windows run START-HERE.bat; on Linux run ./setup.sh.
  3. Accept the recommended answers. The installer picks the model size by RAM and downloads about 70 GB. It resumes if interrupted.
  4. Wait out the first load. Expect a slow or frozen-looking PC for one to three minutes while it loads 35-55 GB.
  5. Connect your tools. Add an "OpenAI-compatible" provider at http://127.0.0.1:8080/v1. For Anthropic-style clients, use /v1/messages (Claude Code reads ANTHROPIC_BASE_URL=http://127.0.0.1:8080). Codex CLI uses /v1/responses.
  6. Set thinking low for speed, high for hard problems.

If you serve it beyond your own machine, the README says always set an API key. For wiring local models into a coding agent, see our guides to running open models locally with OpenCode and fixing local models that loop.

The "let your AI set it up" prompt

Strata's README suggests pasting a one-line prompt into a coding assistant such as Claude Code, Cursor or Codex: set up Strata for me and follow docs/AI_SETUP.md. Hacker News predictably argued about it, comparing it to piping a script into bash. The honest answer is that it is the same trust decision as any installer: read what the agent will run. Commenters did exactly that, reading setup.py and setup.sh before trusting them. The repository also lists the Claude bot among its contributors, and one commenter flagged it, which fits the heavily AI-assisted engineering pattern others noted for inference engines generally.

What this means for what you build or pay

If your cloud coding bill is dominated by well-scoped tasks (refactors, tests, scripted edits) and you already own a 12-24 GB GPU and 64 GB of RAM, Strata is a free experiment worth an afternoon. The upside is privacy, no quota resets to wait for, and a flat hardware cost. The downside is quality you must measure rather than assume. A sensible pattern is routing easy, private jobs to the local model and keeping a hosted model for long-horizon work, the same two-harness hedge we recommend in our Codex 28-day pledge post. For hardware choices, compare against our MacBook vs dedicated GPU guide before buying.

The bigger signal is architectural. Models built on the Qwen4-style sparse design, which Alibaba previewed in Qwen3.8-Flash, reward engines that know where each expert lives. If that holds, expect more model-specific engines and faster consumer-grade local inference.

Related reading

  • Qwen3.8-Flash-Next: the 125B MoE Alibaba teased
  • Qwen3.8-Flash on QwenCloud and the Qwen4 preview
  • OpenCode adds Qwen3.8-Flash 125B
  • Qwen3.8-27B: the local model Hacker News put at #1
  • What is llama.cpp?
  • How to run open models locally with OpenCode
  • MacBook vs dedicated GPU for local LLMs
  • PrismML Bonsai 27B: 1-bit Qwen on a phone

Primary: the Strata GitHub README (Niko1221/Strata, v0.1.39) · the October 5, 2026 Hacker News discussion

Figures are as of October 5, 2026 and come from the project's README and self-reported community posts. Speeds depend on RAM, PCIe lanes, quant and context length; we have not independently benchmarked Strata. Quantized model quality varies by task, so test on your own workload.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 20, 2026

Qwen-Image-2.1: 7B Params, Native Transparency, and a License Downgrade

Qwen-Image-2.1 shrinks Qwen-Image's visual generator from 20B to 7B params, unifies text-to-image and editing with native alpha-channel support, and topped Hacker News at 483 points — but it drops Apache 2.0 for a restrictive new research license Alibaba requires a separate deal to use commercially.

Sep 11, 2026

NVIDIA BioNeMo Inference Runtime: Faster Boltz-2, OpenFold2, Protenix v2

NVIDIA announced the public beta of BioNeMo Inference Runtime on September 10, 2026 — an open-source, PyTorch-native library for speeding up biomolecular structure-prediction inference with specialized kernels, CUDA Graphs, and Ray-powered GPU replica scaling.

Sep 4, 2026

Cerebras Serves Qwen 3.8 27B at 1,500 Tokens/Second — Here's the Catch

Cerebras now serves Alibaba's Qwen 3.8 27B on its wafer-scale hardware at around 1,500 tokens per second, replacing Gemma 4 31B on its shared, pay-as-you-go tier. The raw throughput is real and unmatched by GPU clusters at this price point. What early users are flagging: a 150,000 token-per-minute cap that counts cached tokens at full price, and a 128k context ceiling that together make it expensive for sustained agentic work.