explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What Whistle is
  • How it works
  • Benchmarks, including where Whistle loses
  • How to try it
  • Voice to tool calls in one binary
  • How it compares
  • What this means for what you build or pay
  • Limits and open questions
  • Summary
  • Related reading
← Back to blog

explainx / blog

Cactus Whistle: A 16.9 MB Speech-to-Text Model That Runs on a CPU

Speech to Text, On-Device AI, Open Source AI, Cactus Compute, Edge AI

Part of Local AI

Cactus Compute Whistle is a 16.9 MB open ASR model for seven languages with 11 ms to first token on CPU. WER vs Whisper base, how to run it, and where it loses.

Oct 9, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Cactus Whistle: A 16.9 MB Speech-to-Text Model That Runs on a CPU

Most speech-to-text models are too big for the devices that need them most: watches, earbuds, robots, car dashboards and cheap boards. On October 2, 2026, Cactus Compute released Whistle, a speech recognition model that fits in 16.9 MB and runs on a CPU.

The release is not only a small model. Whistle shares an engine with Cactus's Needle model, so one binary can turn a voice clip into tool calls. This post covers how Whistle works, the benchmark results including where it loses, how to run it, and how it compares with other transcription options on explainx.ai. We read the Hugging Face model card, the Cactus announcement and the Needle 3 card. We did not run Whistle, so the numbers are Cactus's own.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?A 16.9 MB open speech-to-text model for seven languages
License?Apache 2.0
Does it need a GPU?No. CPU only, no dependencies
Speed?11.1 ms to first token and 1,319 tokens per second on an Apple M4 Pro CPU, per Cactus
Better than Whisper base?Wins on some benchmarks, loses on TED-LIUM, AMI and MLS
Languages?English, German, French, Spanish, Italian, Dutch, Polish
Longest clip?30 seconds per pass
Can it act on speech?Yes, with Needle 3 in one call

What Whistle is

The model card describes "a speech-to-text model for mobiles, wearables, robots, smart home, automotive and microcontrollers." The whole model is "a single 16.9 MB file, and it runs on the same CPU engine as Needle, from the same container and the same quantisation, with no dependencies and no GPU."

It does three jobs on the device:

  1. Transcription. 16 kHz mono audio, up to 30 seconds in one pass, in seven languages. The language is detected unless you name it. Silence returns an empty transcript.
  2. Word timestamps. Every word with a start, an end and a probability, aligned from the decoder's own attention. An app can highlight, seek or cut on a word.
  3. Speech embedding. The encoder output, one row per 80 ms frame, for matching and retrieval without decoding a transcript.

It also supports keyword biasing. You pass names, places and product words, and the search favors them so the ones that matter survive.

Whistle at a glance showing a 16.9 MB on-device speech-to-text model with seven languages and 11.1 ms to first token

Figure: Whistle at a glance. Source: the Whistle model card by Cactus Compute (Apache 2.0), used with credit.

How it works

The Cactus announcement, written by Jakub Mroz and Henry Ndubuaku, gives the architecture in detail.

table · 2 cols
StageDetail
Front end16 kHz mono audio, 25 ms window, 10 ms hop, 80 log-mel bins, band-limited to 250 to 3500 Hz
StemA convolutional stem of 128 channels and kernel 9 halves the frame count three times, leaving 375 frames for 30 seconds, one per 80 ms
EncoderEight Simple Attention blocks with four mHC residual lanes and a Monarch Hadamard MLP. Attention is not causal
DecoderEight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, engram lookups at layers 3 and 7 over 18,432 slots
Cross attentionA gated cross attention at each decoder layer reads the encoder. K and V are computed once per clip
DecodingFive beams scored by length-normalized log probability. Cap of 320 tokens. Vocabulary of 8,192 text pieces plus seven language tokens
Quantization2 to 4 bit with Cactus Quants

Three design choices matter for developers.

One engine for speech and text. Whistle reuses Needle's .cact container, Cactus Quants, SIMD kernels and KV cache. A device that already runs Needle runs Whistle in the same binary, with no second runtime.

A decoder ladder. Every decoder depth from 2 layers up was trained as a model of its own. You pick one at load time with --audio-depth. The encoder is never sliced: all eight blocks run at every depth. That gives you a knob to trade accuracy for speed on a weak chip.

Cheap beams. Cross-attention keys and values are computed once when the clip arrives and then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.

Benchmarks, including where Whistle loses

Cactus compares Whistle with Whisper base and Moonshine tiny v2. Word error rate (WER) is scored with the Whisper normalizers. Whistle's WER is measured over 86,174 utterances. The Whisper and Moonshine figures are the ones their authors published, from the multilingual checkpoints.

Whistle word error rate and speed against Whisper base and Moonshine tiny v2 on LibriSpeech, SPGISpeech, Earnings-22, AMI, TED-LIUM, FLEURS and MLS

Figure: Whistle against Whisper base and Moonshine tiny v2. Source: Cactus Compute, from the Whistle model card, used with credit.

Whistle's own WER values in the chart:

table · 2 cols
BenchmarkWhistle WER (%)
LibriSpeech test-clean4.31
LibriSpeech test-other10.49
SPGISpeech7.65
Earnings-2219.01
AMI26.07
AMI cleaned22.87
TED-LIUM7.61
FLEURS (average of seven languages)21.4
MLS (average of six languages, no English)24.9

In its own words, Cactus says: "Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9."

That is a fair reading of the chart. Whistle is not a clean win. It loses on meeting speech (AMI), on TED talks and on the multilingual MLS set. The Whisper figure for AMI is AMI-IHM, a different subset from the AMI the other two report, so that row is not directly comparable.

Speed and size

On 10 seconds of audio on an Apple M4 Pro CPU, with each model on its official runtime at its defaults:

table · 4 cols
MetricWhistleWhisper baseMoonshine tiny v2
Size16.9 MB145.3 MB41.9 MB
Time to first token11.1 ms73.2 ms22.8 ms
Decode speed1,319 tokens/s266 tokens/s262 tokens/s

Whistle ran at 5 beams, Whisper through openai-whisper, and Moonshine through moonshine-voice in non-streaming mode over whole audio. Cactus notes that Whisper pads every input to 30 seconds, so its time to first token stays flat across clip lengths. Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10 seconds and 36.3 ms at 30 seconds.

Read the speed numbers with care. They are for one Apple laptop CPU. A phone, a watch or a microcontroller will run slower, and the card gives no figures for those. Precision also differs: Whistle runs at 2 to 4 bit, Whisper base in fp32 in memory, and Moonshine in int8.

The model card states that no test audio appears in Whistle's training or validation data. Cactus says it checked this by comparing audio checksums and speaker IDs across every reported test set.

How to try it

Install the Python package:

bash
pip install cactus-needle

Then transcribe a clip:

python
import needle

print(needle.transcribe("clip.wav")["text"])
# turn off the kitchen lights

Every call returns the text, the language, the milliseconds to the first token and the decoder's tokens per second. Useful options:

table · 2 cols
OptionEffect
word_timestamps=TrueAdds each word with start, end and probability
keywords=[...]Favors names and product terms during search
language="de"Forces the language instead of detecting it
needle.Whistle()The model as an object for embed(audio) or a tuned .cact

The [mic] extra adds microphone capture and other sample rates. A 16 kHz WAV or raw samples need nothing beyond the base install. The command needle whistle playground transcribes from the microphone in the terminal, and needle whistle compare runs one clip through Whistle, Whisper and Moonshine side by side with timings.

The command-line engine

bash
needle download macos-arm64
needle download whistle
./macos-arm64/needle --model whistle.cact --audio clip.wav --audio-word-timestamps

Per the announcement, the engine ships prebuilt for seventeen targets, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. The speech C API is three functions: needle_load, needle_transcribe and needle_embed. The engine reads no environment variables, so every run of the same file scores the same.

The announcement page also hosts a browser sandbox. The first press downloads the 16.9 MB model, and Cactus says audio never leaves your device.

Voice to tool calls in one binary

The part that sets Whistle apart is the pairing with Needle 3. The Needle card calls it "a foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers," in a single 8 to 29 MB file. It picks functions from your tool list and fills arguments.

Load both models:

bash
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav

The engine transcribes the clip, answers the transcript against your tools and returns one JSON object. A response looks like this:

json
{ "function_calls": [{ "name": "set_lights", "arguments": { "room": "kitchen", "on": false } }],
  "confidence": 0.94,
  "audio_text": "turn off the kitchen lights",
  "audio_language": "en" }

The caller never handles the transcript. For a smart-home panel or a robot, that is one fewer hop and one fewer place for latency and bugs. The confidence field lets you route on act, confirm or refuse.

Edge sensor model on a small device, the kind of hardware an on-device speech-to-text model like Whistle targets

How it compares

table · 3 cols
OptionWhere it runsNotes
WhistleCPU, on device, 16.9 MBSeven languages, 30 s clips, pairs with Needle for tool calls
Whisper baseCPU or GPU, 145.3 MBBroader language coverage, ahead on TED-LIUM, AMI and MLS in Cactus's chart
Moonshine tiny v2CPU, 41.9 MBEnglish only
Superwhisper S1-miniOn deviceA 0.6B cleaner for transcripts, not an ASR model
MAI-Transcribe-2-StreamingCloud APIReal-time streaming, ranked on Artificial Analysis WER
Gemini 3.5 TranscribeCloud APIHosted, large-model accuracy
Meta Muse Voice TranscribeCloudReal-time ASR with diarization and endpointing
Interfaze-1-LiteGPU classOpen weights for OCR, speech and extraction together

The choice is about where audio may go. A cloud model gives higher accuracy on hard audio and many languages. Whistle gives privacy, no per-minute fee, offline use and a tiny footprint. For a voice assistant on a Raspberry Pi or a wearable, those traits decide the choice. For call-center transcription of noisy multi-speaker audio, they do not.

What this means for what you build or pay

  1. Per-minute fees can drop to zero for short commands. Voice control, dictation of short notes and wake-then-command flows fit in 30-second clips. Local inference removes the API bill and the network round trip.
  2. Privacy is easier to promise. Audio stays on the device. The Cactus sandbox states this for its own demo, and the same holds for any app that runs the model locally.
  3. Plan for the weak spots. If your users speak in meetings, on TED-style stages or in languages beyond the seven, test hard before shipping. Whistle trails Whisper base on those sets in the vendor's chart.
  4. One engine simplifies a voice agent. Speech and tool calling share a runtime, a file format and a quantization, so a device carries one binary instead of two stacks.

For a deeper map of building voice agents with open tools, see Hugging Face's speech-to-speech guide.

Limits and open questions

  • Clip length. 30 seconds per pass. Longer audio needs chunking, and the card does not describe a streaming mode.
  • Languages. Seven only.
  • Vendor-measured WER. Whistle's results are Cactus's own, while the baselines come from their papers, which use different test conditions in places.
  • Device coverage. The speed figures are for an Apple M4 Pro CPU. We found no numbers for phones or wearables.
  • Download counts. The model page showed about 2,600 downloads and 196 likes on October 9, so community testing is early.
  • Hacker News. The Cactus post appeared on Hacker News on October 2 with no comments when we checked, so we have no developer reports to quote.
  • Repository access. The Cactus-Compute/needle3 repository, which holds the engine, returned a 404 on GitHub when we fetched it. The Hugging Face repository of the same name loads fine, so we relied on the model cards and the announcement.

Summary

Whistle is a credible tiny ASR model with a clear niche: private, offline, low-latency transcription of short clips on weak hardware, and a one-binary route from voice to tool calls. It is not a Whisper base replacement everywhere. Test it on your own audio, especially meeting speech and non-English languages, before you rely on it.

Related reading

  • MAI-Transcribe-2-Streaming and the Artificial Analysis WER ranking
  • Superwhisper S1-mini: on-device transcript cleaning
  • Gemini 3.5 Transcribe: what builders get
  • Meta Muse Voice Transcribe
  • Hugging Face speech-to-speech voice agent guide
  • Interfaze-1-Lite open weights for OCR and speech

Primary sources: the Whistle model card, the Cactus announcement and the Needle 3 model card.

Specs, benchmark figures and download counts are accurate as of October 9, 2026. Re-check the model card before you build on it.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 10, 2026

Edge0 Open-Sources a 35B Model That Runs in 2.5GB of Memory

Edge0, founded by Samuel Zeng, open-sourced a framework that runs a 35-billion-parameter language model using only 1-2.5GB of peak memory — small enough for a phone, in principle. The technique is "SSD expert offload": stream only the parameters a given step needs from storage instead of loading the whole model into RAM. explainx.ai covers how it works, what's actually shipped today versus the demo, and what it means for on-device AI.

Aug 24, 2026

Pipette: Liquid AI’s Open On-Device Benchmark Suite

On August 24, 2026, Liquid AI and Artificial Analysis released Pipette — an open platform measuring on-device AI as a full system (model × quant × runtime × device), not model cards alone. 1,000+ configs, iOS/Android clients, public data.

Aug 17, 2026

Matic Cues: How a Home Robot Runs Voice, Vision, and Mapping On-Device

Matic Robots launched Cues, a voice-and-gesture control layer for its $115M-funded home robot, built entirely on an Nvidia Jetson Orin Nano. It's a real case study in edge AI system design — wake-word detection, 3D spill localization, and house-scale navigation, all running locally with no data leaving the device.