Most transcription errors are decided before the software ever runs: a noisy room, a distant microphone or clipped audio will trip up even the best speech recognition model. The biggest gains come from recording cleaner audio, preparing the file sensibly, telling the model the language and vocabulary to expect, and then spending ten minutes on a targeted review rather than re-reading everything.
This guide walks through each stage in the order you meet it, with the commands and settings we use when building caption tools.
Why accuracy slips
Speech recognition models, including Whisper, are trained on huge amounts of real-world audio, so they cope with a surprising range of accents and conditions. They still struggle with a predictable set of problems:
- Low signal-to-noise ratio. Air conditioning, traffic, music beds and room echo all compete with the voice.
- Overlapping speech. Two people talking at once usually produces a merged or missing line.
- Clipping and heavy compression. Distorted peaks remove information the model needs.
- Rare words. Brand names, product codes, people's names and niche jargon are often replaced with a common word that sounds similar.
- Long silences and music-only sections. Whisper-style models can "hallucinate" text during stretches with no speech, sometimes repeating a previous line.
If you know which of these you are fighting, you know which fix to reach for. For background on why the model behaves this way, see how Whisper works.
Record for the transcript, not just the edit
Microphone placement
Distance matters more than price. A modest lavalier or dynamic microphone close to the mouth will usually transcribe better than an expensive condenser across the room, because the voice is louder relative to the room.
- Keep a handheld or desk dynamic mic roughly a hand's width from the mouth.
- Clip a lavalier mid-chest, away from clothing that rustles.
- Avoid relying on the camera's built-in microphone unless the camera is very close.
The room
Soft furnishings absorb reflections. Recording in a carpeted room with curtains, a sofa and a bookshelf is noticeably better than a bare kitchen. Switch off fans, fridges and notifications where you can.
Levels
Aim for speech that peaks comfortably below full scale. A common working target is peaks somewhere around -12 to -6 dBFS, leaving headroom for laughs and emphasis. Clipped audio cannot be repaired properly afterwards, whereas slightly quiet audio can be raised.
Separate tracks for separate people
For interviews and podcasts, record each speaker on their own microphone and track where possible. You can transcribe each track independently and merge the results, which avoids most overlap errors.
Prepare the audio file
You do not need a full audio restoration suite. A few FFmpeg steps handle most cases. The >FFmpeg documentation lists every filter used below.
Extract a clean mono file
Whisper resamples input to 16 kHz mono internally, so there is no benefit in sending a large multichannel file. Extracting the audio first also makes uploads and re-runs quicker:
ffmpeg -i interview.mp4 -vn -ac 1 -ar 16000 interview.wav
If the voice is only on one channel of a stereo file (common with some camera setups), pick that channel rather than mixing in a silent or noisy one:
ffmpeg -i interview.mp4 -vn -af "pan=mono|c0=c0" -ar 16000 left-only.wav
Gentle clean-up
A high-pass filter removes low rumble, and a light denoiser can help with steady hiss:
ffmpeg -i interview.wav -af "highpass=f=80,afftdn=nf=-25" interview-clean.wav
Be restrained. Aggressive noise reduction creates watery artefacts that can make recognition worse, not better. Always compare a short section of the raw and cleaned files with the same settings before committing.
Even out the loudness
If one speaker is much quieter than another, loudness normalisation helps both come through:
ffmpeg -i interview-clean.wav -af loudnorm -ar 16000 interview-norm.wav
The -ar 16000 matters: loudnorm works internally at 192 kHz, and without it the output file is written at that rate.
Trim dead air
Long stretches of silence or music at the start and end are where hallucinated lines tend to appear. Trim them, or at least note their timings so you can check those regions first.
Choose sensible model settings
The examples below use the open-source Whisper command-line tool from the >official repository. Other Whisper-based tools expose similar options under slightly different names.
Pick the right model size
Larger models are generally more accurate but slower and need more memory. For English speech on clean audio, a mid-sized model is often good enough; for accents, crosstalk or other languages, the larger models earn their keep. Our guide to choosing a Whisper model size covers the trade-offs in detail.
Set the language
Language auto-detection uses the first part of the audio. If your video opens with music or a non-English clip, detection can go wrong and the whole transcript suffers. Set it explicitly:
whisper interview-norm.wav --model medium --language en
Prime it with vocabulary
The --initial_prompt option gives the model some text to condition on before it starts. It is a hint rather than a guarantee, but it is useful for names and spellings:
whisper interview-norm.wav --model medium --language en \
--initial_prompt "Guests: Siobhan Ní Bhriain, Tomasz Kowalczyk. Topics: Premiere Pro, DaVinci Resolve, LUTs, B-roll."
Writing the prompt in the style you want back (proper capitalisation, British spellings, punctuation) also nudges the output towards that style.
Reduce repetition loops
If you see the same sentence repeated over and over, the model has latched on to its own previous output. Turning off conditioning on previous text often breaks the loop, at the cost of slightly less consistent style between segments:
whisper interview-norm.wav --model medium --language en --condition_on_previous_text False
A worked example
Suppose a 20-minute interview was recorded on a camera microphone in an office, with a guest whose surname keeps being mis-heard. A practical sequence is:
- Extract mono 16 kHz audio and listen to thirty seconds. The air conditioning hum is obvious.
- Apply
highpass=f=80and a lightafftdn. Listen again at the same volume: the hum is lower, the voice is not warbly. - Run a mid-sized model with
--language enand an initial prompt containing the guest's name and company. - Search the transcript for the guest's name and a couple of technical terms. Fix any stragglers with find-and-replace.
- Jump to the intro music and the last minute of the file, where invented lines are most likely, and check them against the audio.
That targeted routine usually catches the bulk of meaningful errors far faster than a line-by-line proofread.
Review efficiently
Even a very good transcript needs a human pass before it becomes published captions.
- Search for known risk words: names, numbers, brands and anything in your initial prompt.
- Check numbers and units. "Fifteen" and "fifty" are easily confused, as are "£1.5 million" and "£15 million".
- Scan for repeated lines and text in sections you know are silent.
- Listen at 1.25× to 1.5× speed with the transcript scrolling alongside. You will notice mismatches quickly.
- Keep a house style list (e.g. "YouTube", "TikTok", "colour grade") and apply it consistently.
Once the words are right, check the result as captions: line length and reading speed matter too. Our caption reading speed guide and the caption speed checker help there.
Common mistakes
| Mistake | What happens | Fix |
|---|---|---|
| Relying on auto language detection | Whole file transcribed in the wrong language or translated | Set --language explicitly |
| Over-processing noise | Robotic voice, more errors | Use light filters and compare |
| Leaving long silences in | Invented or repeated lines | Trim, or check those regions first |
| Clipped recording | Unrecoverable distortion | Record with headroom |
| Proofreading everything equally | Slow, and tired eyes miss the real errors | Target names, numbers and silent sections |
Quick checklist
- Microphone close to the speaker, room as quiet and soft as possible
- Peaks well below 0 dBFS, no clipping
- Separate tracks for separate speakers where you can
- Extract mono 16 kHz audio; apply only gentle clean-up
- Set the language explicitly
- Add names and jargon to an initial prompt
- Choose a model size that suits the audio difficulty
- Review names, numbers, repeated lines and silent sections first
If privacy is a concern for the audio you are transcribing, read our comparison of local vs cloud transcription before choosing a service.
FAQ
Does a more expensive microphone always give better transcripts?
No. Placement and room noise usually matter more than price. A basic microphone close to the speaker in a quiet room often beats a premium microphone placed far away.
Should I run noise reduction before transcribing?
Only lightly. A high-pass filter and mild denoising can help with steady noise, but heavy processing introduces artefacts that can increase errors. Test on a short clip first.
Why does Whisper sometimes add sentences nobody said?
Long silences, music and very noisy sections give the model little to work with, and it can generate plausible-sounding text or repeat earlier lines. Trimming those sections and checking them during review deals with most cases.
Can the initial prompt force the correct spelling of a name?
It improves the odds but does not guarantee it. Always search the finished transcript for important names and correct any that slipped through.