Skip to content

Build1 publisher3 min readPublished

Cactus Compute fits a 16.9MB speech model into the CPU runtime it built for tool calling

Cactus Compute released Whistle, an open 16.9MB speech-to-text model that runs on a CPU in the same Needle engine as its tool-calling language models. The accuracy and speed figures are the company's own, so they need rerunning on real devices and audio.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Cactus Compute fits a 16.9MB speech model into the CPU runtime it built for tool calling
Generated illustration

What happened

  • Cactus says a combined setup can pass an audio clip straight to a language model and get structured tool calls back, extending its Needle 2 work on low-cost devices.
  • Whistle covers seven languages and transcribes up to 30 seconds of 16 kHz mono audio in one pass, returning word-level timestamps and probabilities.
  • According to Cactus, low-volume silence and steady noise produce an empty transcript instead of a guessed sentence.
  • Cactus reports word error rates of 4.31% on LibriSpeech test-clean and 10.49% on test-other, against 4.9% and 11.0% for Whisper base.
  • On an Apple M4 Pro CPU, Cactus measured 11.1 ms to first token and 1,319 decoded tokens per second, against 73.2 ms and 266 for Whisper base.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability A team already on Needle can take spoken input to a structured device action on the device itself, shipping one engine and one model format.
  • exposure Once transcripts drive tool calls, a sentence guessed from room noise becomes an action the user never asked for. Runtimewire says the launch material does not establish how reliably the features hold up across accents and noise.
  • decision Choosing between Whistle and Whisper base turns on whether a product's audio resembles the three sets Whistle wins or the three it loses. Only a test on in-house recordings settles it.
  • constraint Apps that capture more than 30 seconds at a time, such as dictation, need to split audio before each pass or confirm the runtime does it for them.

Whistle ships in Needle's .cact model container and runs on Needle's CPU inference engine [7]. A team already deploying one of Cactus's tool-calling models gets speech input without taking on a second runtime or model format [7]. According to Cactus, the engine builds for 17 platform targets, including RISC-V, MIPS, WebAssembly and WASI [16].

The model is an encoder-decoder [5]. A log-mel front end and a convolutional stem feed an eight-block audio encoder, and a decoder with gated cross-attention writes the text [5]. The company says the decoder can be loaded at different depths while the encoder always runs at full depth [6]. Trimming the decoder makes each token cheaper, but every clip still pays for the whole encoder pass [6]. The launch material does not say which decoder depth produced the published figures, or how long a ten-second clip takes from first sample to finished transcript.

On Cactus's figures, Whistle reaches its first token about 6.6 times sooner than Whisper base and decodes about five times faster [1][2]. The speed test used ten seconds of audio and reported time to first token separately from decoding [15]. I think reporting the two apart is correct for a model of this shape. The first number includes the encoder pass, and the second is decoder work alone.

Needle 2, the release Whistle builds on, was about running tool-calling models on low-cost devices [8]. For the speed ratios to carry over to cheaper hardware, both models would have to slow down by the same factor on a weaker CPU. The published timings cover one Apple M4 Pro, against 17 build targets [14][16].

Accuracy splits three sets each way among the head-to-head comparisons Cactus names [6]. Whistle takes both LibriSpeech sets and FLEURS, while the company's own benchmark page gives TED-LIUM, AMI and the MLS average to Whisper base [11]. The LibriSpeech margins are 0.59 and 0.51 points of word error rate [3][4]. A half-point lead transfers to a product only if its audio resembles LibriSpeech more than the sets Whisper base wins. On FLEURS the gap is wider, at 3.1 points [5].

The test design shows care. According to Cactus, its word-error runs span 86,174 utterances, and the company compared the test sets with the data Whistle was trained and validated on [13]. That contamination check is the first thing I'd ask any vendor about. Comparison models ran on their official runtimes with default settings, and published Whisper and Moonshine results were used where available [12]. Runtimewire's report notes that the figures remain company-reported [18].

For a voice-command feature on hardware a team already targets with Needle, I'd bench Whistle before Whisper base. Adding it means one more model in a container format the app already loads [7]. For recordings that look like TED-LIUM, AMI or MLS, Cactus's own table favours Whisper base [11].

What to watch

  • An independent reproduction of Whistle's LibriSpeech and FLEURS word error rates, or of the M4 Pro speed figures.
  • Published latency for Whistle on the low-cost and mobile targets that Needle 2 was built for.
  • Measurements of the clip-to-tool-call pipeline end to end, including how often silence or noise triggers a spurious call.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories