Word-level timestamps give every individual word in a transcript its own start and end time, rather than timing whole sentences. You can get them in two main ways: ask a speech recognition model such as Whisper to estimate them while it transcribes, or run forced alignment, which takes a transcript you already trust and lines each word up against the audio. The first is quicker; the second is usually more precise.
Why word timings matter
Most caption files are timed at the segment level: a block of text appears at one time and disappears at another. That is fine for traditional subtitles, but a lot of modern creator work needs finer control:
- Short-form captions that show two to four words at a time, or highlight each word as it is spoken (see our guide to short-form captions).
- Cutting to the word when editing talking-head footage, so you can remove an "erm" without clipping the next word.
- Syncing visuals to narration, for example changing an image exactly when the narrator says the name of a place.
- Re-segmenting captions to meet a reading-speed target, which is far easier when you know where every word falls.
Segment timings force you to guess where inside a sentence a word sits. Word timings remove the guesswork.
Two approaches: estimated versus aligned
It helps to be clear about what each method is actually doing.
Timings estimated during transcription
When you ask Whisper for word timestamps, it transcribes as normal and then estimates where each word falls. OpenAI's implementation does this by looking at the model's cross-attention (which parts of the audio the decoder was "looking at" when it produced each token) and applying dynamic time warping to turn that into a path through time. If you want the background, our explainer on how Whisper works covers the encoder–decoder structure this relies on.
The advantage is that you get text and timings in one pass. The drawback is that the timings are a by-product of recognition, not the goal of it. They are often good enough for word-by-word captions, but you will sometimes see words that start a little early, end a little late, or bunch up after a pause or a burst of music.
Forced alignment
Forced alignment starts from the opposite end. You already have the correct text, perhaps a script, a corrected transcript or a lyric sheet, and the aligner's only job is to decide when each word was spoken. Because it is not also trying to work out what was said, it can concentrate on when.
Aligners typically use an acoustic model that knows what speech sounds (phonemes) look like, convert your text into the expected sequence of sounds, and then find the best path through the audio that matches that sequence. Well-known open-source options include:
| Tool | What it does | Good for |
|---|---|---|
| WhisperX | Transcribes with Whisper, then re-aligns words using a separate phoneme recognition model | Fast, accurate word timings from raw audio |
| Montreal Forced Aligner | Classic forced aligner using pronunciation dictionaries and acoustic models | Known transcripts, research-grade phone and word boundaries |
| Timings from Whisper alone | Cross-attention estimate, no separate aligner | Quick jobs where "close enough" is fine |
The main catch with forced alignment is that the text must match the audio. If the narrator ad-libbed a sentence that is not in your script, the aligner will try to squeeze your text into audio that does not contain it, and timings around that point will go wrong.
Getting word timestamps from Whisper
The reference openai-whisper command-line tool can output word timings directly. A typical command looks like this:
whisper narration.wav --model small --language en --word_timestamps True --output_format json
The JSON output contains a list of segments, and each segment has a words array. A trimmed example:
{
"start": 0.0,
"end": 3.2,
"text": " Welcome back to the workshop.",
"words": [
{ "word": " Welcome", "start": 0.0, "end": 0.42, "probability": 0.97 },
{ "word": " back", "start": 0.42, "end": 0.66, "probability": 0.99 },
{ "word": " to", "start": 0.66, "end": 0.78, "probability": 0.99 },
{ "word": " the", "start": 0.78, "end": 0.9, "probability": 0.99 },
{ "word": " workshop.", "start": 0.9, "end": 1.54, "probability": 0.94 }
]
}
A few things to notice:
- Leading spaces are part of each word. Strip them before building captions, but keep punctuation, which is attached to the word it follows.
- Times are in seconds as decimals, not SRT-style
00:00:01,540strings. You convert them when you write the caption file. probabilityis the model's confidence in that word. Low values are a useful flag for words worth checking by ear.
If you only want the result as captions, the same tool can also write an SRT with word highlighting using --highlight_words True, and limit line length with --max_line_width and --max_line_count. Those options depend on word timestamps being switched on. The full list is in the project's README on github.com/openai/whisper.
If you use the faster-whisper library from Python, the equivalent is to pass word_timestamps=True to transcribe() and read segment.words, where each item has start, end, word and probability.
Model size affects timing quality indirectly: a larger model makes fewer recognition mistakes, and wrong words tend to have wrong timings. Our guide to Whisper model sizes explains the trade-offs.
Turning word timings into captions
Once you have a list of words with start and end times, building captions is a grouping problem. A simple, robust approach:
- Clean the list. Strip leading spaces, and drop any empty tokens.
- Group words into chunks. For short-form, two to four words per chunk. For standard subtitles, group until you reach your line-length limit.
- Break at natural points. Prefer to end a chunk at punctuation or before a conjunction, rather than splitting "the" from its noun.
- Use the first word's start and the last word's end as the chunk's timing.
- Close small gaps. If the next chunk starts within about a quarter of a second, extend the current chunk's end to meet it, so captions do not flicker off and on.
- Enforce a minimum duration. A one-word caption that lasts 0.12 seconds cannot be read. Stretch very short chunks, borrowing time from a following pause where there is one.
- Check reading speed. Run the result through the caption speed checker to catch chunks that are too dense.
Here is a small Python sketch of steps 2 and 4, assuming words is a list of dictionaries with word, start and end:
def chunk_words(words, max_words=3):
chunks, current = [], []
for w in words:
current.append(w)
ends_phrase = w["word"].rstrip().endswith((".", ",", "?", "!"))
if len(current) >= max_words or ends_phrase:
chunks.append(current)
current = []
if current:
chunks.append(current)
return [
{
"text": " ".join(x["word"].strip() for x in c),
"start": c[0]["start"],
"end": c[-1]["end"],
}
for c in chunks
]
This is deliberately minimal. In production you would add the gap-closing and minimum-duration rules above, and handle hyphenated words and numbers that the model splits into several tokens.
Common problems and how to fix them
Words drift late after music or silence. Timings estimated during transcription are most fragile around long pauses and background music. If you see a run of words starting too late after an intro, try trimming leading silence before transcription, or run a forced aligner over the corrected text.
Everything is shifted by the same amount. If every word is consistently early or late, the timings themselves are probably fine and something upstream has offset them, such as a video with an audio delay or a clip that was trimmed after transcription. Shift the whole file rather than re-transcribing; the SRT shifter does this in the browser. If the offset grows over time instead, you have drift rather than an offset, which our guide to fixing subtitle sync drift covers.
The last word of a sentence hangs on too long. Recognition models often stretch the final word into the following pause. Cap the end of any word at, say, 0.6 seconds beyond its start unless the next word begins later than that.
Numbers, names and hyphenated terms split oddly. "2026" may come back as two tokens, and a brand name may be broken in the middle. Merge adjacent tokens with no space between them before grouping.
The transcript was corrected by hand, and now the timings do not match. If you edited words in a text editor, the original timings refer to the old words. This is exactly the case forced alignment is designed for: align the corrected text against the audio again rather than patching times by hand.
Quick checklist
- Decide whether you need speed (Whisper word timings) or precision (forced alignment on a corrected transcript).
- Use a clean, consistent audio file; strip long silences and loud music from the start if you can.
- Turn on word timestamps explicitly; they are not on by default in the Whisper CLI.
- Strip leading spaces and merge split tokens before grouping.
- Group words at natural phrase boundaries, not purely by count.
- Enforce minimum durations and close tiny gaps.
- Spot-check low-probability words and the first few seconds after any pause.
- Check reading speed before you publish.
FAQ
Are Whisper's word timestamps accurate enough for captions?
Usually, yes, for captions that show a few words at a time. They can be less reliable around pauses, music and overlapping speech. For karaoke-style highlighting or tight edits, a forced alignment pass gives cleaner results.
What is the difference between transcription and forced alignment?
Transcription works out what was said and, optionally, when. Forced alignment assumes you already know what was said and only works out when each word was spoken, which is why it needs an accurate transcript as input.
Can I get word timestamps from an existing SRT file?
Not directly. An SRT only stores timings per caption block. You can run a forced aligner using the SRT's text and the original audio to recover word-level timings.
Do word timestamps work for languages other than English?
Whisper can produce word timings for many languages, and aligners such as WhisperX and the Montreal Forced Aligner support a range of languages through language-specific models. Quality varies by language, so test with a short clip before committing to a long job.