Cactus Compute releases a 16.9MB speech model for local CPUs
Co-founder Henry Ndubuaku helped build Whistle into Cactus Compute's existing Needle runtime, adding keyword biasing and word timestamps for on-device applications.
By RuntimeWire Staff · Published
Primary source: Cactus
Why it matters
Whistle extends Cactus's on-device strategy from local tool-calling models into speech. Its benchmarks are company-reported, and performance still needs to be judged on the hardware and audio that real applications use.

Cactus Compute released Whistle on October 2nd, an open speech-to-text model designed to run locally on a CPU. At 16.9MB, it is built to fit the same Needle runtime that Cactus Compute uses for small on-device language models, extending co-founder Henry Ndubuaku's work on compact inference into speech recognition. The release is covered in Cactus Compute's technical launch post.
Cactus Compute's release thread on X
Ndubuaku is Cactus Compute's co-founder and CTO. His work has spanned fundamental AI research, distributed deep learning, GPU-kernel engineering and inference on small devices; before Cactus Compute, he was a research engineer on mobile AI models at an Imperial College spinout, according to his Forbes Technology Council profile. Cactus Compute CEO Roman Shemet, a former quant and economist with product and data-engineering experience, met Ndubuaku through Y Combinator's co-founder-matching program in London, according to YC's company profile. Their bet is that useful AI can move onto the devices people already carry, rather than requiring every interaction to travel to a server.
A speech model built to fit the existing stack
Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. Cactus Compute says it can transcribe a 16 kHz mono recording of up to 30 seconds in one pass, return word-level timestamps and probabilities, and provide speech embeddings without generating a transcript. Developers can also pass a list of keywords to bias recognition toward names or domain-specific terms. Cactus Compute says low-volume silence and steady noise return an empty transcript rather than a guessed sentence.
Keyword biasing could help an application catch a customer's name; word timestamps can support highlighting, seeking or editing; and an empty result on silence addresses a familiar failure mode in voice interfaces. The launch material does not establish how reliably these features perform across real users, accents and noisy environments.

The model uses a log-mel audio front end, a convolutional stem, an eight-block audio encoder and a decoder with gated cross-attention. Cactus Compute says the decoder can be loaded at different depths, while the audio encoder stays at full depth. Whistle also shares Needle's .cact model container and CPU inference engine. In practice, Cactus Compute is adding speech input to a deployment system it already wants developers to use for local tool-calling models.
This builds on Cactus Compute's earlier Needle 2 release, which focused on running tool-calling models on low-cost devices. Whistle can transcribe audio in the same engine, and Cactus Compute says a combined setup can pass a clip directly to a language model and return structured tool calls. That makes voice a route into local device actions as well as a transcription feature.
Cactus Compute's benchmark results
Cactus Compute reports Whistle at 4.31% word error rate on LibriSpeech test-clean and 10.49% on test-other, compared with 4.9% and 11.0% for Whisper base. It reports a 21.4 average on FLEURS versus 24.5 for Whisper base. Cactus Compute also lists results on SPGISpeech and Earnings-22, where it says Whistle scores 7.65 and 19.01 respectively.

Those results do not show Whistle winning across the board. Cactus Compute's own benchmark page says Whisper base performs better on TED-LIUM, AMI and the MLS average. The comparisons use published results for Whisper and Moonshine where available, and Cactus Compute says the models ran on their official runtimes with default settings. It reports that its word-error-rate tests cover 86,174 utterances and that it checked the test sets against Whistle's training and validation data. Those details describe the methodology, but the figures remain company-reported rather than an independent reproduction.
On an Apple M4 Pro CPU, Cactus Compute reports 11.1 milliseconds to first token and 1,319 decoded tokens per second for Whistle, against 73.2 milliseconds and 266 tokens per second for Whisper base. The speed comparison uses ten seconds of audio and separates time to first token from subsequent decoding. Cactus Compute also says its engine ships for 17 platform targets, ranging from desktop and mobile systems to RISC-V, MIPS, WebAssembly and WASI. The reported speed is tied to the M4 Pro test; it does not establish equivalent performance on every listed device.
Why the deployment choice matters
Whistle is aimed at applications where sending audio to a cloud service is undesirable or unavailable: wearables, phones, robots, cars and smart-home devices. Local inference can avoid a network round trip and keep audio on the device, while a compact model can be easier to fit into constrained hardware. Cactus Compute has also described a hybrid approach in which local inference handles routine audio and cloud processing can take more difficult segments. Whistle's local-only design and the broader hybrid product are distinct deployment choices.
Whistle's scope is limited: it supports seven languages and caps a single pass at 30 seconds; the published comparisons show wins over Whisper base on some datasets and losses on others. Teams evaluating it will need to test their own audio, target hardware and vocabulary. Cactus Compute includes a compare command that runs a clip through Whistle, Whisper and Moonshine with timing results, and its launch materials provide a Python install path through cactus-needle.
For Ndubuaku and Shemet, the release extends their engineering thesis: make local models small enough to run on ordinary devices, then build useful product behavior around them. Whistle adds speech to that stack. The next test is whether developers can get dependable transcription on the constrained devices Cactus Compute targets, not just strong scores on selected benchmarks.