Whisper is an open-source speech recognition model from OpenAI that turns audio into text by first converting sound into a picture-like representation called a log-Mel spectrogram, then passing it through a Transformer that "reads" the audio and writes out text one token at a time. It was trained on a very large amount of real-world audio paired with transcripts, which is why it copes well with accents, background noise and many languages without special tuning.
Knowing roughly what happens inside helps you make better decisions: which settings to change, why it sometimes invents text in silent passages, and why its timestamps behave the way they do.
The big picture
Whisper uses an encoder-decoder Transformer, the same broad family of architecture used for machine translation. You can think of it as two halves:
- The encoder listens. It takes 30 seconds of audio and turns it into a rich internal representation of what was said and how.
- The decoder writes. It looks at that representation and produces text, one small piece at a time, each time predicting the most likely next piece given everything so far.
According to the original research paper, the first Whisper models were trained on 680,000 hours of multilingual and multitask audio collected from the web. Rather than training separate systems for each job, OpenAI trained one model to handle several tasks: transcribing speech in its own language, translating speech into English, identifying the spoken language and detecting whether there is any speech at all.
Step 1: Audio becomes a spectrogram
Whisper does not work directly on your MP3 or MP4. The reference implementation calls FFmpeg to decode the file and resample it to 16 kHz mono audio. That is why FFmpeg must be installed before Whisper will run.
The audio is then converted into a log-Mel spectrogram. In plain terms:
- The sound is chopped into very short overlapping slices.
- For each slice, the model measures how much energy there is at different frequencies.
- Those frequencies are grouped onto the Mel scale, which spaces them roughly the way human hearing does, with more detail in the lower ranges where speech carries most of its information.
- The values are put on a logarithmic scale, again closer to how we perceive loudness.
The result looks like a heat map with time running left to right and pitch running bottom to top. Most Whisper models use 80 Mel frequency bands; large-v3 and the turbo model use 128.
Step 2: The 30-second window
Whisper always processes audio in 30-second windows. Shorter clips are padded with silence up to 30 seconds; longer recordings are handled window by window.
For a long file, the transcription code works through the audio sequentially. After decoding one window, it uses the timestamps it predicted to decide where the next window should start, and by default it passes the text it just produced into the next window as context. That context helps keep spelling and style consistent, but it also explains one of Whisper's quirks, which we come back to below.
Step 3: The encoder
The spectrogram first passes through two small convolutional layers, which pick out local patterns and reduce the length of the sequence. Position information is added so the model knows the order of events in time. The result then flows through a stack of Transformer blocks.
Each Transformer block uses self-attention: every moment in the audio can "look at" every other moment in the same window and weigh how relevant it is. That is how the model can use a word later in the window to disambiguate a sound earlier on. Larger Whisper models simply have more of these layers and wider ones, which is where the difference in accuracy and speed comes from. Our guide to Whisper model sizes compares them.
Step 4: The decoder and special tokens
The decoder produces tokens, which are pieces of words (a common word might be one token; an unusual name might be several). It is steered by a short sequence of special tokens at the start, which effectively tell it what job to do:
| Token (simplified) | What it tells the model |
|---|---|
| Start of transcript | Begin a new output |
| Language, e.g. English or Welsh | Which language the speech is in |
| Transcribe or Translate | Write in the original language, or translate into English |
| No timestamps (optional) | Skip timing information |
If you do not specify a language, Whisper first predicts the language token itself from the opening audio, then carries on. Setting it explicitly avoids misdetection on short or noisy clips:
whisper interview.wav --model small --language en --output_format srt
The translate task only ever outputs English. There is no setting to translate into other languages; for that, transcribe first and translate the text separately, as covered in translating subtitles.
Step 5: Timestamps
When timestamps are enabled, the decoder also predicts special timestamp tokens between phrases, marking where a segment starts and ends. These are quantised to 20-millisecond steps relative to the start of the window. That is how Whisper produces segment-level timings for SRT and VTT output.
Segment timestamps are good but not frame-perfect, and they describe phrases rather than individual words. The reference implementation offers a --word_timestamps True option, which estimates word timings using the model's attention patterns and a technique called dynamic time warping. It is useful, but for precise work many people run a separate forced-alignment step. Word-level timestamps explains the options.
Step 6: Choosing the words
At each step the decoder has a probability for every possible next token. Depending on settings, Whisper either takes the single most likely token each time (greedy decoding) or keeps several candidate sentences alive and picks the best (beam search; the command-line tool uses a beam size of 5 by default). Either way there is a safety net: if the output looks wrong, it tries again with more randomness.
"Looks wrong" is measured two ways:
- Compression ratio. If the text compresses too well, it is probably repeating itself.
- Average log probability. If the model was not confident across the segment, the result is suspect.
When either threshold is crossed, Whisper re-decodes that window at a higher temperature, which lets it pick less likely tokens and often escapes a repetition loop. A separate check uses the "no speech" probability to skip windows that appear to be silence.
Why Whisper sometimes hallucinates
Because the decoder is essentially a language model trained to produce plausible text, it will sometimes produce plausible text when there is no speech to transcribe. Common symptoms include:
- A phrase repeated over and over during music or silence.
- Invented sign-offs such as a "thanks for watching" line at the end of a clip, which are common in web video transcripts the model learned from.
- Text copied from the previous window, carried over by the context mechanism.
Practical fixes:
- Trim long silences and music-only sections before transcribing.
- Try
--condition_on_previous_text Falseif one error keeps propagating through a long file. - Use a voice activity detection (VAD) step to feed Whisper only the parts with speech; several community implementations include this.
- Review the transcript at every point where the audio is quiet.
Our transcription accuracy tips cover recording and preparation in more depth.
Guiding it with a prompt
The --initial_prompt option lets you pass text that the model treats as if it came just before the audio. It is not an instruction, more a hint about style and vocabulary. Including correctly spelt names and jargon often improves how they are written:
whisper episode12.mp3 --model medium --language en --initial_prompt "Guests: Siobhan Ní Bhriain, Aneurin Pryce. Topics: DaVinci Resolve, LUTs, Rec.709."
Keep the prompt short and relevant. A long or misleading prompt can make results worse.
Key takeaways
- Whisper converts audio to a log-Mel spectrogram, then an encoder-decoder Transformer turns it into text tokens.
- It always works in 30-second windows and carries text context between them by default.
- Special tokens select the language and the task; the translate task only outputs English.
- Timestamps are predicted as tokens at 20 ms resolution; word timings are an extra estimation step.
- Hallucinations come from its language-model nature; trimming silence and using VAD help most.
The source code, model list and command-line options are documented in the >openai/whisper repository.
FAQ
Does Whisper need an internet connection?
No. Once the model files have been downloaded, the open-source version runs entirely on your own computer. Whether audio leaves your machine depends on the tool you use, not on Whisper itself; see local vs cloud transcription.
Why does Whisper need FFmpeg?
The reference implementation uses FFmpeg to read almost any audio or video format and convert it to 16 kHz mono before building the spectrogram.
Can Whisper tell who is speaking?
Not on its own. Whisper transcribes speech but does not label speakers. Speaker diarisation requires a separate tool, whose output is then combined with the transcript.
Is Whisper the same as the OpenAI transcription API?
They are related but not identical. The open-source models run locally; OpenAI's hosted API offers transcription models on its servers under its own terms. Check OpenAI's current documentation for which models the API uses.