Which Whisper Model Size Should You Use? Tiny to Large-v3 and Turbo

Compare Whisper's tiny, base, small, medium, large and turbo models, plus the English-only .en versions, and pick the right one for your hardware and audio.

Speech to TextBy AI Point EditorialUpdated 24 September 20267 min read

For most creators transcribing clear English speech on a reasonably modern computer, turbo or small.en is the sensible starting point: turbo if you have a graphics card with around 6 GB of memory, small.en if you do not. Move up to large-v3 when accuracy on difficult audio or non-English languages matters more than speed, and only drop to tiny or base for quick drafts or very limited hardware.

This guide explains what each model is, how the English-only versions differ, and how to test which one suits your audio.

The official model list

The figures below come from the model table in the >openai/whisper README. Memory and speed figures are OpenAI's own approximate numbers: speed is relative to the large model on an A100 GPU, so treat them as a rough ranking rather than a promise for your machine.

Model Parameters English-only version Approx. VRAM Approx. relative speed
tiny 39 M tiny.en ~1 GB ~10x
base 74 M base.en ~1 GB ~7x
small 244 M small.en ~2 GB ~4x
medium 769 M medium.en ~5 GB ~2x
large 1,550 M none ~10 GB 1x
turbo 809 M none ~6 GB ~8x

"Parameters" is the number of learned values in the model. More parameters generally means better accuracy and slower, more memory-hungry transcription.

Large, large-v2 and large-v3

"Large" is not a single model. There have been three releases:

  • large (sometimes called large-v1): the original large model released with Whisper in 2022.
  • large-v2: a retrained version with the same architecture, trained for longer with extra regularisation. It improved accuracy over the original.
  • large-v3: a later release with the same basic size, trained on considerably more data. It uses 128 Mel frequency bands for its audio input instead of 80, and added a language token for Cantonese.

In the reference Python package, asking for large loads the newest large model, which at the time of writing is large-v3. You can request a specific version by name:

whisper talk.mp3 --model large-v2
whisper talk.mp3 --model large-v3

Is large-v3 always better than large-v2? Usually, but not universally. Some users have found v2 behaves better on particular audio, for example with fewer repetitions in long, quiet recordings. If you rely on the large model for important work, test both on a sample of your own material.

What turbo is

Turbo (also referred to as large-v3-turbo) is an optimised version of large-v3. OpenAI cut the decoder from 32 layers to 4, which makes it much faster, while keeping the full large-v3 encoder. The README describes the accuracy loss as minimal.

The important limitation: turbo was not trained for the translation task. If you run it with --task translate, it will generally return the original-language text rather than an English translation. For translation into English, use medium or large instead.

For plain transcription, turbo is often the best balance of speed and accuracy available in the official set, provided your hardware has enough memory for it.

The English-only .en models

The tiny, base, small and medium models each come in two flavours: multilingual (for example small) and English-only (small.en). The English-only versions were trained just for English transcription.

According to the README, the .en models tend to perform better for English, particularly tiny.en and base.en. The advantage becomes less significant at small.en and medium.en. There is no large.en or turbo.en.

When to use them:

  • Your audio is entirely in English: use the .en version at the small sizes, and test it against the multilingual one at medium.
  • Your audio contains any other language, or you need the translate task: use the multilingual model. The .en models cannot identify languages or translate.
whisper podcast.m4a --model small.en --output_format srt

How to choose: a practical decision path

  1. Check your hardware. Do you have an NVIDIA GPU, and how much video memory? On Windows, Task Manager's Performance tab shows "Dedicated GPU memory". Without a supported GPU, the reference implementation runs on the CPU, which is much slower; the larger models may be impractically slow on long files.
  2. Check your audio and language. Clear, single-speaker English in a quiet room is easy for any model. Accents, crosstalk, noise, music beds, jargon and non-English speech all favour larger models.
  3. Decide what you need the output for. A rough transcript for finding quotes in an interview can come from a fast model. Published captions need the best accuracy you can reasonably get, followed by a human review.
  4. Test on your own material. Pick a representative two- to three-minute clip, including a difficult section, and run it through two or three candidate models.

A quick comparison loop on Windows PowerShell:

foreach ($m in "small.en","medium.en","turbo") {
  whisper sample.wav --model $m --language en --output_format txt --output_dir "test_$m"
}

Or in a Bash shell:

for m in small.en medium.en turbo; do
  whisper sample.wav --model "$m" --language en --output_format txt --output_dir "test_$m"
done

Then compare the text files side by side. Count the errors that would actually need fixing: misheard names, missing words, invented phrases. Also note how long each run took. The right model is usually the smallest one whose error count you are happy to correct by hand.

Suggested starting points

Situation Try first Step up to
Clear English, modest laptop, no GPU base.en or small.en medium.en if time allows
Clear English, GPU with 6 GB or more turbo large-v3
Accents, noise or several speakers turbo or medium large-v3
Non-English or mixed-language audio turbo or medium large-v3
Foreign audio to English subtitles medium large-v3 (not turbo)
Very quick rough draft tiny.en or base.en small.en

These are starting points, not verdicts. Your microphone, room and speakers matter as much as the model; our transcription accuracy tips often gain more than a model upgrade.

Things that matter as much as model size

  • Set the language. --language en avoids a misdetected language on short or noisy clips, which can wreck an otherwise good transcript.
  • Trim silence and music. All sizes can hallucinate text in long quiet sections. Larger models are not immune.
  • Use a vocabulary hint. --initial_prompt with correctly spelt names helps every model.
  • Consider faster implementations. Community projects such as faster-whisper and whisper.cpp run the same model weights with lower memory use or better CPU speed. Results can differ slightly from the reference package, so test before switching.
  • Plan for timing work. Model size affects text accuracy more than timing accuracy. If you need precise word timings for animated captions, read about word-level timestamps.

Key takeaways

  • Official sizes: tiny, base, small, medium, large and turbo; tiny to medium also have English-only .en versions.
  • The .en models help most at tiny and base; the gap narrows at small and medium.
  • Large has three versions; large currently loads large-v3 in the reference package.
  • Turbo is a faster large-v3 with a 4-layer decoder, and it does not translate.
  • Test two or three candidates on your own audio and pick the smallest acceptable one.

FAQ

Is the biggest model always the most accurate?

Generally the larger models make fewer errors, especially on difficult audio, but the improvement on clean English can be small. On some recordings a smaller model even avoids a repetition that a larger one makes. Test on your own material.

Can I use turbo to translate a Spanish video into English subtitles?

Not reliably. Turbo was not trained for translation and tends to return Spanish text. Use medium or large-v3 with --task translate, or transcribe in Spanish and translate the text separately.

Where are the model files stored?

The reference package downloads each model the first time you use it and caches it in your user profile (by default a .cache/whisper folder). You can choose another location with --model_dir.

Does a bigger model give better timestamps?

Not necessarily. Segment timings come from the same mechanism across all sizes and remain approximate. For precise timings, add a forced-alignment step, and always check sync in a player before publishing. For how the model works internally, see how Whisper works.

Software, platform rules and settings change. We review our guides regularly, but always check the official documentation for the tools you use. Found an error? Email soubickdas@gmail.com. See our editorial policy.
A

AI Point Editorial

We build caption, transcription and video-workflow tools and write about what we learn doing it — practical, tested and free of hype.