Speech runs at 120–150 words a minute. Good typing runs at 60–80. So dictation should be roughly twice as fast - and for most people it isn't, because they dictate the way they type: a sentence at a time, watching the screen, correcting as they go.
The techniques below are in rough order of how much difference they make. The first three account for most of the gain.
This is the single biggest change. Watching words materialise pulls you into editing mode, and you can't compose and edit at the same time - you stop mid-sentence to fix a comma and lose the thread.
Look away from the screen. Look out of a window. Some people close their eyes. Say the whole paragraph, then look.
If your app makes this hard, that's an argument for a tool that inserts text in one block rather than streaming word by word. Most local apps - superwhisper, VoiceInk, Handy - do exactly that: you hold a key, talk, release, and the finished text lands.
Speech recognition uses surrounding context to disambiguate. "There," "their," and "they're" are identical acoustically; the model picks one based on what came before and after. Feed it three words and it's guessing. Feed it a full sentence and it usually gets it right.
Longer utterances are also faster: fewer starts and stops, less latency overhead, and better punctuation because the model can see the sentence shape.
Aim for at least a full sentence per push of the key. A paragraph is better.
Dictate the whole thing badly, then fix it. Do not correct as you go.
Correcting mid-flow costs more than the errors do: you break composition, you switch input modes, and you often re-say things you'd already said. Getting 800 rough words down in six minutes and spending four minutes editing beats 800 careful words in twenty.
This is also the argument for tools that clean up after you. Wispr Flow strips filler words and false starts before the text appears. Amical does the same locally, free, pairing Whisper with an Ollama model. Both reduce how much editing the draft needs.
The biggest single accuracy variable is not the software. A $60 USB microphone or a decent headset will do more for your error rate than switching from a good app to a slightly better one.
Laptop microphone arrays are tuned for calls: noise suppression, automatic gain, aggressive processing that helps intelligibility on Zoom and hurts transcription. A microphone six inches from your mouth, off-axis so plosives miss it, in a room with soft furnishings, is a different signal entirely.
See the best microphone for dictation.
If you say the same fifteen proper nouns every day and correct all fifteen every day, you're paying a tax you could remove in ten minutes.
Add them to the custom dictionary. Wispr Flow learns terms after a few corrections and supports shared team dictionaries. superwhisper, VoiceInk, and Voice In all take explicit vocabulary lists. Dragon Professional still does this better than anything else.
Apple Dictation doesn't offer it, which is the main reason heavy users move off it - see Apple Dictation vs paid apps.
You need "new paragraph" and "new line." Beyond that, saying "comma" and "period" constantly is slow and breaks your rhythm.
Modern tools punctuate automatically and well - Apple Dictation added on-device automatic punctuation in 2026, and every AI-cleanup tool does it. Let them. Speak with natural sentence intonation and pauses, and check the punctuation during the editing pass.
Most dictation apps expand a trigger phrase into a block of text. Email sign-offs, standard disclaimers, addresses, common technical phrases, a client's full legal name.
Wispr Flow has snippets built in. superwhisper modes can go further and trigger shell commands.
The way you want a Slack message formatted is not the way you want an email formatted, which is not the way you want a commit message formatted. Switching manually is friction; switching automatically isn't.
VoiceInk switches modes per app automatically at $29 one-time. superwhisper has custom modes. Wispr Flow adapts tone to the app without configuration.
Reverb hurts recognition more than steady background noise does. A hard-surfaced room with a bare desk and a window is a bad recording environment regardless of your microphone.
Cheap wins: close the window, turn off the fan, put something soft on the desk, avoid sitting in a corner. If you dictate in an open-plan office, a headset with a boom arm is close to mandatory.
Bigger Whisper models are more accurate and slower. On an M-series Mac you can usually run a large model in real time. On an older CPU you can't, and a laggy large model will slow you down more than a fast small one costs you in errors.
Try the tiers. See Whisper models explained.
Speech is bad at structure. Left to talk freely, most people produce fluent paragraphs that don't build an argument.
Dictate the headings first, then fill each section. You get the speed of speech and the structure of writing, and it removes the most common complaint about dictated prose - that it rambles.
Dictation feels slower for about a fortnight. You're learning to compose out loud, which is a genuinely different skill from composing while typing, and it's uncomfortable at first.
Most people who quit do so in week one. Most people who get past week three don't go back.
Start where the stakes are low - Slack messages, notes to yourself, first drafts nobody sees - and expand from there.
Speed. 120–150 words a minute of raw output is achievable. Net of editing, most people land somewhere between 1.5× and 2× their typing speed for first drafts.
Where it doesn't help. Anything structurally dense: tables, code, heavily formatted documents, or text you're already composing slowly because it's hard to think through. Dictation speeds up transcription of thought, not thinking.
Your voice will tire. Two or three hours a day is a lot. Alternate with typing.
Accuracy plateaus. Once you have a decent microphone, a quiet room, and your vocabulary loaded, further gains come from your technique rather than the software.
Speech runs 120–150 words a minute against 60–80 for good typing. Net of editing, most people get 1.5–2× their typing speed on first drafts.
In order of likelihood: your microphone, your room, missing custom vocabulary, speaking in fragments rather than sentences, and only then the software. Fix them in that order.
Only "new paragraph" and "new line." Modern tools punctuate sentences automatically and well, and saying every comma slows you down and disrupts your rhythm.
For first drafts, yes, once you've practised. For editing, formatting, and structurally dense text, no - typing wins.
About three weeks to feel natural. It's slower than typing for the first week or two while you learn to compose out loud.
Low latency matters most: Aqua Voice is the fastest cloud option, and local apps like superwhisper and VoiceInk have no network round trip at all. See the best dictation software.
Last verified: 10 August 2026.