People shopping for better dictation accuracy almost always compare software. It's the wrong variable. The signal reaching the model is set by your microphone, its position, and your room - and improving those three usually beats the difference between any two good apps.
If your dictation is inaccurate and you're using a laptop's built-in microphone, buy a microphone before you buy software.
A laptop's microphone array is engineered for video calls, and the engineering actively works against transcription.
It's far away. Two to three feet from your mouth, picking up as much room as voice.
It processes aggressively. Noise suppression, automatic gain control, and beam-forming are tuned to make you intelligible to a human on a compressed call. Speech models were trained on relatively clean audio and don't benefit from the same processing - sometimes it actively harms them by chopping consonants.
It hears the machine. Fan noise, keyboard, and desk vibration all reach it.
Move to a microphone six inches from your mouth and the signal-to-noise ratio improves dramatically before any software is involved.
Under $50 - a wired headset. Any decent headset with a boom microphone that sits near the corner of your mouth. Consistent distance regardless of how you move, and cheap. This is the best value in the entire category and what most heavy dictation users end up with.
$60–120 - a USB condenser or dynamic microphone. A desk microphone on a stand or arm, positioned six to eight inches away and slightly off-axis. Better sound quality, no headset to wear, but distance varies as you shift in your chair.
Dynamic microphones reject room noise better than condensers, which matters more than raw fidelity here. You are not recording a podcast; you want a clean close signal.
$150+ - an XLR microphone and interface. Better, but well past diminishing returns for transcription. Buy this if you also record audio for other people to listen to.
Wireless earbuds - avoid for dictation. AirPods and similar switch to a low-bandwidth codec when the microphone is active, and the resulting audio is noticeably worse than the same earbuds' playback quality suggests. Fine for a quick note; not for sustained work.
Distance: 15–20cm (6–8 inches). Closer picks up breath and plosives; further picks up the room.
Off-axis. Point the microphone at the corner of your mouth rather than straight at it. This alone removes most plosive damage from "p" and "b" sounds.
Consistency. Speech models handle a consistently imperfect signal better than a varying one. A boom arm or headset keeps the distance fixed.
Isolate it from the desk. Typing and desk knocks transmit through the stand. A shock mount, or even a folded towel, helps.
Check your input level. Aim for peaks around -12 to -6 dB. Clipping destroys accuracy; too quiet forces the software to amplify noise along with your voice.
Reverb hurts recognition more than steady background noise. A voice bouncing off hard surfaces arrives at the microphone smeared, and smeared consonants are what models get wrong.
In order of effect and cost:
None of this costs anything, and together it can matter as much as the microphone.
Operating systems and conferencing apps apply enhancements designed for calls. For dictation, they usually hurt.
macOS: turn off Voice Isolation in Control Centre while dictating. It's tuned for calls.
Windows: in Sound settings, disable audio enhancements and any manufacturer "AI noise cancellation" on your input device.
Conferencing apps: Zoom, Teams, and similar can hold the microphone and apply their own processing. Close them.
The exception is Krisp, which does on-device, real-time, bidirectional noise cancellation and is genuinely good - but it's designed for calls. Test with and without it on your own dictation before leaving it on.
Yes, once the signal is good. But the order of operations matters:
If you've done all five and it's still poor, the problem may be accent coverage in the model, which is a known and unevenly distributed weakness. Larger models help. See speech-to-text accuracy explained.
Steps three and four are also where the free tools stop being a compromise: superwhisper, VoiceInk, and Handy all let you load a vocabulary list and pick a larger model, and Handy costs nothing. Browse the full directory to filter by platform and price.
Record the same three-minute passage under each configuration and compare transcripts against what you actually said.
Worth testing: built-in microphone versus headset; headset versus desk microphone; enhancements on versus off; quiet room versus normal room; six inches versus eighteen. Change one variable at a time.
Most people find the built-in-to-headset step is the biggest single improvement they'll make, and it costs less than two months of a dictation subscription.
Different constraints, and worth planning for.
Walking or driving. Wired earbuds with an inline microphone beat wireless. Whisper Memos at $60 a year is built for this - talk into an Apple Watch or a lock screen widget and get an email.
Phone dictation. Hold the phone normally rather than at arm's length; the bottom microphone is designed for a call position. Apple Dictation runs on-device and handles this well.
Open-plan offices. A headset with a boom and good rejection is close to mandatory, both for accuracy and for the people around you.
A wired headset with a boom microphone is the best value - consistent distance, good rejection, and usually under $50. A USB dynamic microphone on an arm at six to eight inches is the next step up.
More than the software choice. Moving from a laptop's built-in array to a close headset typically improves error rate more than switching between two good dictation apps.
Not really. Wireless earbuds drop to a low-bandwidth codec when the microphone is active, so the recorded audio is worse than their playback quality suggests. Fine for a quick note, poor for sustained work.
Usually yes. Voice Isolation on macOS and audio enhancements on Windows are tuned for calls and can remove detail the speech model uses. Test both ways on your own voice.
15–20cm (6–8 inches), positioned slightly off-axis so plosives don't hit it directly. Consistency matters as much as the exact distance.
It helps, because a cleaner signal always helps, but accent coverage is a model limitation rather than a signal one. Running a larger model - see Whisper models explained - makes more difference there.
Last verified: 10 August 2026.