Whistle packs speech recognition into 16.9 MB and transcribes seven languages on-device

Whistle: Speech to Text in 16.9 MB

Cactus Compute released Whistle, a 16.9 MB speech recognition model that runs on CPU with no dependencies and shares the same C++ engine as Needle. It transcribes English, German, French, Spanish, Italian, Dutch, and Polish, reaches the first token in 11 ms, and handles word timestamps and speech embeddings. Benchmarks show lower word error rates than Whisper base and Moonshine tiny v2 on several sets, while being far smaller and faster. It also integrates with Needle for direct speech-to-tool-calls.

Whistle does three jobs, all of them on the device: transcription, word timestamps, and speech embedding.
  1. skolos

    Interesting that this is here. I used whistle (and bunch of other things) to take ownership of my echo show. It now doesn't dial to Amazon at all - it does all processing locally with its own CPU and connects to my homeassistant for home automation. My initial setup involved qwen asr (1.7b model) running on rtx 5080. Compared to that, whistle was really bad (out of 170 messages, qwen recognized correctly 168, whistle - 70), but I adjusted whistle to work like jev - instead of free form transcription it recognizes only select set of templates (I trained tiny network with 10,000 generated utterances to translate whistle final state to probabilities within templates). The precision went up to 164/170 - almost matching qwen. By the way - I'm speaking with heavy accent.

  2. INTPenis

    I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.

    I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.

    But every single sound he makes with his mouth ends up on the page too.

  3. albert_e

    What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.

  4. zimpenfish

    Tried it on a random TV episode and it seems to get stuck sometimes where it just outputs "Thank you." as a default - at one point emitting that for 60s of dialogue (and no, the episode does not have someone repeating "Thank you." for 60s.) Happens several times during the transcription.

  5. wkcheng

    How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)

    This definitely seems lighter and faster. How does accuracy compare?

  6. flowerlad

    Apple needs to incorporate this into iOS ASAP. This works much better than the speech recognition in iOS when you use technical terms. One of the most frustrating parts of iOS is speech-to-text in iMessage. For me no feature is more important in a phone.

    Try this example: My website uses ASP.NET technology and I am using .NET 10.0. Works perfectly in Whistle, but not in iOS.

  7. andy_ppp

    Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!

    I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!

  8. joewhale

    I initially read this as whistle to text, which would be way cooler.

More from this day

2026-10-08