Skip to main content
Back to Blog
Blog

Local, Private Voice Input with On-Device Whisper

Why MadoHub runs speech recognition on your machine instead of calling a cloud API, what that means for privacy and cost, and how to get the most out of dictation.

MadoHub sits next to your codebase. When you dictate into it, you're often describing the thing you're currently looking at — a stack trace, a variable name, a snippet of a config file, sometimes a secret you're reading off a .env file to explain to an agent. That's not audio you want to hand to a third-party speech API by default. So voice input in MadoHub doesn't call one. It runs whisper.cpp locally, in the same process as the rest of the app.

This post is about why that decision matters and how to get the most out of dictation in MadoHub, at a level of detail you can act on.

How MadoHub handles it

No audio leaves the machine

There's no streaming upload, no API key, no request going out to a speech-to-text vendor. The captured audio is transcribed by a local Whisper model that lives on your disk. The only network call in the entire feature is the one-time model download; after that, recording and transcription both happen with the machine offline.

That has two direct consequences worth naming:

  • Nothing about what you say (or what's on your screen while you say it) is transmitted anywhere. For a tool built to sit inside your working codebase, that's not a nice-to-have — it's the reason voice input is usable for anything sensitive at all.
  • There's no per-minute or per-request cost. Once a model is downloaded, transcribing is free and unmetered — it's local compute, not a billed API call.

Works offline

Because the model lives on disk and the transcriber runs in-process, voice input works with no internet connection at all, as long as the model file is already present. MadoHub checks this before it even opens the microphone: if the selected model isn't downloaded, it sends you to Settings → Voice Input with a prompt to download it, rather than recording into silence and only failing later at transcription time.

Padded for short utterances, filtered for noise

Whisper's encoder was trained on 30-second chunks and copes better with a couple of seconds of context. Very short utterances — a quick word or two — can trigger a known Whisper failure mode where the decoder loops the same phrase over and over. MadoHub guards against this by padding short recordings with a brief moment of silence before transcription, giving the decoder enough room to find a coherent single pass.

Quiet or ambiguous audio can still produce garbage, so MadoHub applies two thresholds to drop segments that look like silence or low-confidence noise before they become text. The result is that quiet stretches don't turn into hallucinated transcripts in your prompt.

Loaded once, reused, and unloaded on demand

The Whisper model is lazy-loaded: nothing is read off disk until the first transcription request. Once loaded, it stays resident in memory and is reused across calls — for the default small model, that's roughly 500 MB held in RAM for as long as voice input is in use. If you switch model size in Settings → Voice Input, MadoHub drops the old model rather than letting a stale buffer linger, and loads the newly selected one on the next transcription.

Practical guidance

  • Download a model before you need it. Voice input won't open the microphone until a model is present. Pick one in Settings → Voice Input ahead of time so the first dictation isn't blocked on a download.
  • Choose a model size that fits your machine. The default small model is a good balance of accuracy and footprint (~500 MB resident). Larger models are more accurate but heavier; smaller ones are lighter but less accurate. Switch in Settings → Voice Input and the change takes effect on the next transcription.
  • Speak naturally, not in single words. Because Whisper handles 30-second chunks better than fractions of a second, normal sentences transcribe more reliably than isolated keywords. If you dictate one word at a time, expect occasional repetition artifacts.
  • Use it for what's on your screen. Dictation shines when you're describing a stack trace, walking through a bug, or narrating a plan while looking at code — exactly the cases where you'd rather not hand the audio to a cloud API.
  • Stay offline if you want. Once the model is downloaded, transcription works with the network entirely off. If you're working on something sensitive, that's a meaningful property, not a curiosity.

None of this is exotic engineering — the appeal is that the entire path from microphone to text is short, local, and inspectable, which is exactly what you want when the thing you're dictating might be a description of code that isn't supposed to leave your laptop.