Voice Input
Dictate into MadoHub with on-device speech-to-text — your audio never leaves the machine.
Voice input lets you dictate instead of type. Hold a hotkey, speak, release — the transcript lands wherever you've told it to go. The whole pipeline runs locally: audio capture, resampling, and transcription all happen on your machine, with no network call in between.
If you have ever dictated into a cloud speech API and wondered where that audio ended up, this is the other design. Your voice is captured, transcribed by a local Whisper model, and written into MadoHub — all in the same process, on your machine. Nothing about the audio or the transcript is uploaded.
What you can do with it
- Push-to-talk dictation — hold a hotkey, speak, release to transcribe.
- Retry the last recording — re-run transcription on the same captured audio without speaking again, useful when the first pass misheard you.
- Route the transcript to one of three destinations — into the MadoAgent chat input, into the focused terminal without submitting, or into the focused terminal and submit immediately.
- Switch destinations on the fly — change where dictation lands from the command palette without opening Settings.
- Pick a model size — from
tiny(~75 MB) up tomedium(~1.5 GB); larger models are more accurate but slower and heavier to download. - Pin a language or auto-detect — English, Chinese, Japanese, Korean, Spanish, French, German, or auto.
How to use it
- Enable voice input from Settings → Voice, or from the command palette (Cmd+Shift+P) via the Enable Voice Input entry (searchable by "voice", "dictation", "push-to-talk", or "语音").
- If you have not downloaded a Whisper model yet, MadoHub opens Settings on the Voice tab with a toast telling you what to download — it will not record into silence.
- Hold Cmd+Shift+V (the default push-to-talk hotkey) to record, and release to stop and transcribe.
- Press Cmd+Shift+R to retry the last recording if the first pass misheard you. Retry only works within 30 seconds of a successful transcription and only if you have not started a new recording since.
- Choose where transcripts land in Settings → Voice, or switch destinations on the fly from the command palette.
| Destination | Behavior |
|---|---|
| Agent (default) | Drafted into the MadoAgent chat input, which auto-opens so you see it land. You review, edit, and press Enter yourself. |
| Terminal — insert | Written into the currently focused terminal tile without submitting. Newlines collapse to spaces so it does not auto-submit in a shell. |
| Terminal — auto | Written into the focused terminal tile and submitted immediately. |
The two terminal destinations lock to whichever terminal tile was focused when you started recording — a focus change mid-recording cannot misroute the transcript. If no terminal was focused (or the tile closed before you finished speaking), MadoHub falls back to the Agent destination and shows a toast explaining why.
Common use cases
- "Refactor the auth module to use a class, add tests for the login flow." — dictate a long prompt into MadoAgent while walking through the code in your head, then review and send.
- Running a quick shell command without typing — set the destination to Terminal — auto and dictate
git statusorpnpm test. - CJK dictation — pin the language to Chinese/Japanese/Korean for better accuracy on short utterances, since Whisper's auto-detect can misfire on quiet or very short CJK clips.
Tips and best practices
- Start with the
smallmodel (the default) for a good accuracy/size tradeoff. Move up tomediumonly if you dictate often and need the extra accuracy. - Pin a language when you know what you are speaking — it avoids auto-detect mistakes on short or quiet clips.
- Rebind the push-to-talk and retry hotkeys in Settings if they collide with shortcuts you already use; Settings warns you about conflicts.
- Use Terminal — insert rather than Terminal — auto when dictating multi-line input, so you can review before submitting.
Related features
- Command Palette — toggle voice input and switch destinations from here.
- Settings — the Voice tab where you configure model, language, hotkeys, and destination.
- MadoAgent — the default dictation destination.
Limitations
- Voice input ships off by default; enable it before first use.
- Retry only works within 30 seconds of a successful transcription, and only if you have not started a new recording since.
- Terminal destinations require a live terminal tile focused at the moment you start recording. Otherwise MadoHub falls back to the Agent destination.
- A model must be downloaded before recording; MadoHub will not record into silence.