explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: questions people will ask first
  • What exactly was released
  • How to run it
  • The benchmark table, read carefully
  • Where it loses
  • Who should use it
  • How it compares with other extreme-compression efforts
  • What to check before you rely on it
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

Underdog Saluki 27B: Qwen3.8-27B Squeezed Into 7.89 GB With Tool Calling Intact

Open Source, Local AI, Qwen, Quantization, Tool Calling

Part of AI Tools and Apps

Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B at 7.89 GB that scores 88 vs 84 on tool calling. How to run it in llama.cpp and what it gives up.

Oct 10, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Underdog Saluki 27B: Qwen3.8-27B Squeezed Into 7.89 GB With Tool Calling Intact

Underdog Saluki 27B 1.0 is a 2-bit GGUF build of Qwen3.8-27B that fits in a single 7.89 GB file, down from roughly 54 GB for the full-size model. Its author, ConwayResearch, published it on Hugging Face under Apache 2.0 and says it was tuned to keep tool calling intact. On the card's own 120-task tool-calling test, it passes 88 tasks against 84 for the uncompressed original.

That is a surprising result for a model squeezed to about a seventh of its size, and it is worth reading with some care. This post covers what the model is, how to run it with stock llama.cpp, what the benchmark table does and does not prove, and who should reach for it.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: questions people will ask first

table · 2 cols
QuestionAnswer
What is it?Qwen3.8-27B compressed to a 2-bit "IQ2-mix" GGUF, 7.89 GB
Who made it?ConwayResearch (the card credits "Underdog"), built on the Qwen team's model and ISTA-DASLab's quantized base
License?Apache 2.0, same as both bases
Does it need special software?No. Stock llama.cpp and apps built on it
Does it do images?Yes, with an optional 629 MB or 928 MB vision add-on file
Is it faster?Unknown. The card publishes no tokens-per-second figures
Does it beat the full model?On tool calling, by the card's test. On math and long reasoning, no
Downloads so far?15,274 in the last month and 184 likes at the time of writing

What exactly was released

The repository holds one main file, Underdog-Saluki-27B-1.0-IQ2-mix.gguf, plus two optional vision projector files. The model card describes the base as Qwen3.8-27B and names ISTA-DASLab's Qwen3.8-27B-GSQ-RCO-GGUF as the quantized derivative it builds on. The architecture tag is qwen35 and the model is 27 billion parameters.

The card is vague on method. "IQ2-mix" signals a mixed 2-bit scheme in the llama.cpp family of quantization formats, and the repository tags include imatrix, which usually means an importance matrix was used to decide which weights deserve more precision. But the card does not describe the mix, the calibration data, or what the "tuned to keep tool calling intact" step actually involved. That is the main gap for anyone who wants to reproduce the result.

For background on the base model family, see our coverage of Qwen3.8 Flash and the Qwen 4 architecture preview and the Qwen3.8 Flash Next 125B MoE release. The 27B size is the one hardware vendors keep showing off: Cerebras ran Qwen3.8 27B at about 1,500 tokens per second, and the Oki Home memory computer ships with a local copy.

How to run it

The model card gives two commands. First, download the file:

bash
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir .

Then serve it with llama.cpp:

bash
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768

The --jinja flag matters. It turns on the Qwen3.8 chat template, which is what handles tool calls and the thinking mode. Without it you lose both. The flags -ngl 99 offload all layers to the GPU, -fa on enables flash attention, and -c 32768 sets a 32K context window. The server speaks the OpenAI chat API on port 8080, so most agent frameworks can point at it with a base URL change.

To add vision, download one of the mmproj files (928 MB in F16 or 629 MB in Q8_0) and pass it with --mmproj. The main file is text only.

The card recommends two settings profiles:

table · 2 cols
UseSettings
General use, reasoning, instructionsThinking on (default), temperature 0.6, top_p 0.95, top_k 20
Fast, direct tool callsThinking off via chat_template_kwargs with enable_thinking false, temperature 0

That second profile is how the tool-calling numbers below were produced, so use it if you want to match them.

The benchmark table, read carefully

A gripper choosing the right tool from a tray, illustrating tool calling in a small local modelA gripper choosing the right tool from a tray, illustrating tool calling in a small local model

The headline result is on "Underdog Bench," which the card describes as 120 tasks taken from the Berkeley Function Calling Leaderboard (BFCL v4), frozen before any model was tested, run with thinking off at temperature 0.

table · 3 cols
ModelSizePassed (of 120)
Underdog Saluki 27B 1.07.89 GB88
Qwen3.8-27B, full size54 GB84
Bonsai 25.95 GB70

On 100 BFCL v4 parallel-call tasks scored with the official checker, Saluki gets 42 and the full model 35. These two results are the basis for the claim that tool calling survived compression.

Two cautions. The card itself says 120 tasks is a modest test and that a few tasks of difference is within run-to-run variation, so 88 against 84 is better read as "roughly on par" than "better." And these are the authors' own numbers on a set they selected; nobody independent has reproduced them yet. Still, parity on tool calling at one seventh of the size is the claim that matters for agent builders, and it is easy to test yourself.

Where it loses

The wider table is more sobering, and the card deserves credit for publishing it.

table · 3 cols
BenchmarkSalukiFull size
SWE-bench Verified (50 issues fixed)3033
IFEval (prompt-loose)93.591.5 (public)
IFBench (prompt-loose)72.771.0 (public)
MBPP+78.083.9 (public)
MuSR67.579.6 (public)
AIME 2025 (avg@4)79.296.7 (public)
AIME 2026 (avg@4)80.094.6 (public)

A test bench with a probe touching a block, illustrating benchmark evaluation of a quantized modelA test bench with a probe touching a block, illustrating benchmark evaluation of a quantized model

Instruction following holds up, coding drops a little, and competition math falls by around 15 points. The card puts the retention at about 82 to 85 percent of the full model's math score and says average retention across nine benchmarks is 96 percent. It also notes that the "public" scores for the full model come from a different harness than the one used for Saluki, so the comparison is not apples to apples; where the authors ran the full model themselves, they used the same harness and settings.

The card lists further limitations: it is weakest at letter-level puzzles such as palindromes and alphabetical ordering, about one in five parallel-call replies carry small formatting slips, and with thinking on it often reasons at length before answering. For a 2-bit model, none of this is shocking. Letter-level tasks are a known weak spot for tokenized models, and aggressive quantization tends to cost the most on long multi-step reasoning, which is exactly what AIME measures.

Who should use it

The sensible use case is a local agent that spends most of its time choosing and calling tools rather than proving theorems: a coding helper wired to a few commands, a home-automation controller, a document pipeline, or a private assistant that must stay on the machine. Our guide to agent skills and the explainer on MCP show the kinds of tool surfaces such a model would drive.

It is a poor pick if you need strong math, long chains of reasoning, or exact string manipulation. In that case a larger quantization of the same base, or the full model, is still the safer choice.

Memory is the other consideration. A file under 8 GB brings a 27B-class model within reach of a 12 GB graphics card or a laptop with 16 GB of unified memory, though the KV cache for a 32K context adds to that, and the card gives no minimum hardware spec. If you are weighing machines for this kind of workload, our Surface Laptop Ultra configuration guide for local AI walks through how much memory different model sizes need, and llama.cpp's decision model support covers recent runtime changes worth keeping current with.

How it compares with other extreme-compression efforts

Saluki sits in a crowded field of "small enough to run anywhere" work. Samsung's LittleBit pushes a 13B model under 1 GB with sub-1-bit weights, but it is a research paper result rather than a file you can load in llama.cpp today. Saluki's pitch is the opposite: a modest 2-bit level, a standard GGUF, and a benchmark focused on the one capability agents depend on. If you have been following the on-device privacy push, Underdog's private personal AI launch is a separate product post of ours; the model card does not say the two are connected, so we are not claiming they are.

What to check before you rely on it

  1. Run your own tool schema. BFCL tasks are not your tools. Test with your actual function definitions, including nested arguments and parallel calls.
  2. Watch for formatting slips. The card admits roughly 20 percent of parallel replies have small formatting errors. Add schema validation and a retry.
  3. Pick thinking mode deliberately. Off and temperature 0 for fast tool calls; on for harder reasoning, accepting the longer outputs.
  4. Measure speed. With no published tokens per second, benchmark on your hardware before building a latency budget around it.
  5. Check the NOTICE file. The card points to a NOTICE file for attribution to the Qwen team and ISTA-DASLab, both Apache 2.0.

Bottom line

Underdog Saluki 27B is a practical, openly licensed shrink of a strong base model, with an honest limitations section and a claim that is cheap to verify. Treat the tool-calling win as promising rather than proven, treat the math losses as real, and download it if you want a 27B-class agent brain in under 8 GB. The unanswered questions are how the "tuned to keep tool calling intact" step works and whether independent testers reproduce the 88 and 42 scores.

Model details, file sizes and benchmark figures are from the Hugging Face model card as of October 10, 2026 and may change as the repository is updated.

Related reading

  • Qwen3.8 Flash and the Qwen 4 architecture preview
  • Qwen3.8 Flash Next 125B MoE release
  • Cerebras runs Qwen3.8 27B at 1,500 tokens per second
  • Oki Home: a local Qwen 3.8 27B memory computer
  • Samsung LittleBit: a 13B LLM under 1 GB
  • llama.cpp decision model support
  • What is MCP
  • llama.cpp releases and the base Qwen3.8-27B model
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Aug 20, 2026

Unsloth Ships Dynamic v3.0 GGUFs for Qwen3.8-27B — What Quant Should You Run?

Unsloth followed up its early-preview Dynamic v3.0 quantization method with a full release for Qwen3.8-27B — GGUFs it says beat every other provider's quants by more than 10% top-1% accuracy at matched size, down to a 1-bit build that still holds 77% accuracy on 8GB of RAM. This is the practical guide to which quant fits your hardware and how to set it up.

Oct 9, 2026

Qwen-Image-2.1-Turbo: The 8-Step 7B Image Model, and How to Run It

Qwen has published Qwen-Image-2.1-Turbo, an accelerated checkpoint of its 7B Qwen-Image-2.1 that generates and edits images in 8 denoising steps with CFG set to 1. Here is how to run it, what the settings mean, and what the license allows.

Oct 9, 2026

Samsung LittleBit: A 13B LLM Under 1 GB, and What the Paper Really Shows

Samsung Research released LittleBit code that compresses LLMs to 0.1 bits per weight by factorizing and binarizing each weight matrix. The headline says a 13B model fits under 1 GB. The paper says so too, but the method needs training, and the license blocks commercial use.