How to Translate Subtitles Without Breaking the Timing

A practical workflow for translating SRT and VTT subtitles: lock the source timing, translate text only, then re-check reading speed and line breaks.

Captions & SubtitlesBy AI Point EditorialUpdated 24 September 20268 min read

The safest way to translate subtitles is to finish and lock the timing of your source-language file first, then translate only the text inside each cue, leaving every index number and timecode untouched. Afterwards you check each translated cue for reading speed and line length, because the same sentence can be noticeably longer in German or French than in English, and a cue that was comfortable in the original may flash past too quickly in translation.

This guide covers the workflow we use for SRT and WebVTT files, the problems that most often break timing, and how to fix them without re-timing the whole video.

Why translation breaks subtitle timing

A subtitle file is two things stitched together: a timing track (when each cue appears and disappears) and a text track (what the cue says). Timing problems after translation almost always come from one of these mistakes:

  • The timecodes were edited by accident. Some translation tools, and some people working in a word processor, change --> into an arrow character, swap commas for full stops in the milliseconds, or "smart quote" the file. Players then reject or mis-read the cue.
  • Cues were merged or split during translation. A translator rewrites two short English cues as one longer sentence, and now there is one block of text sitting across two timing slots, or an empty cue.
  • The translated text is longer than the time allows. The cue still appears at the right moment, but viewers cannot finish reading it before it disappears.
  • The source file was never properly synced. If the original drifts, every translation inherits the drift. Fix that first; our guide on fixing subtitle sync drift explains how.

The first two are process problems and are easy to prevent. The third is a genuine translation problem, and it is where most of your review time should go.

Step 1: Finalise the source file

Before anyone translates a word, make the source subtitles as good as they will ever be:

  1. Correct transcription errors, names and spellings.
  2. Check sync against the final cut of the video, not an earlier edit.
  3. Split long cues so each one carries a single idea. Cues that already break at natural clause boundaries translate far more cleanly.
  4. Check reading speed in the source language. If a cue is already fast in English, it will be too fast in almost any translation.

If the edit changes after translation starts, you will be patching timing in every language separately. Lock the picture, lock the source subtitles, then translate.

Step 2: Separate the text from the timing

The most reliable approach is to hand translators, or a translation tool, only the text, with a stable ID for each cue. That way nothing can touch the timecodes.

A minimal SRT cue looks like this (see what an SRT file is for the full structure):

12
00:00:41,200 --> 00:00:44,050
We tested three microphones
in the same room.

A short Python script can pull the text into a tab-separated sheet and later put translated text back into the original structure. This version uses only the standard library:

import re, csv, sys

def read_srt(path):
    text = open(path, encoding="utf-8-sig").read().replace("\r\n", "\n")
    blocks = [b for b in text.strip().split("\n\n") if b.strip()]
    cues = []
    for b in blocks:
        lines = b.split("\n")
        cues.append({"id": lines[0].strip(), "time": lines[1].strip(),
                     "text": "\n".join(lines[2:])})
    return cues

def export(srt_path, tsv_path):
    with open(tsv_path, "w", encoding="utf-8", newline="") as f:
        w = csv.writer(f, delimiter="\t")
        w.writerow(["id", "source", "translation"])
        for c in read_srt(srt_path):
            w.writerow([c["id"], c["text"].replace("\n", " | "), ""])

def rebuild(srt_path, tsv_path, out_path):
    cues = read_srt(srt_path)
    with open(tsv_path, encoding="utf-8") as f:
        rows = {r["id"]: r["translation"] for r in csv.DictReader(f, delimiter="\t")}
    with open(out_path, "w", encoding="utf-8") as f:
        for c in cues:
            t = rows.get(c["id"], "").strip() or c["text"]
            f.write(f'{c["id"]}\n{c["time"]}\n{t.replace(" | ", chr(10))}\n\n')

if __name__ == "__main__":
    if sys.argv[1] == "export":
        export(sys.argv[2], sys.argv[3])
    else:
        rebuild(sys.argv[2], sys.argv[3], sys.argv[4])

Usage:

python subs.py export video.en.srt video.tsv
python subs.py rebuild video.en.srt video.tsv video.de.srt

The | marker preserves the original line break so the translator can see where it was. Any cue left blank in the sheet falls back to the source text, which makes gaps obvious when you review. The timing line is copied straight from the source, so it cannot change.

Step 3: Translate cue by cue, not sentence by sentence

This is where human translators and machine translation both need guidance. Subtitle translation is not the same as document translation:

  • One cue in, one cue out. If a sentence runs across two cues, the translation must also be split across the same two cues, even if the word order has to change. Languages with verb-final clauses, such as German, make this harder, so allow the translator to rephrase rather than translate literally.
  • Condense, do not pad. Subtitles routinely shorten speech. Dropping filler words, repetitions and "you know" is normal and expected.
  • Keep on-screen references aligned. If the speaker says "this one" while pointing, the translated cue must still appear while they point.
  • Keep names, brands and numbers consistent. Give translators a short glossary of terms that must not be translated or must always be translated the same way.

If you use machine translation, feed it the whole script for context but insist on the cue-by-cue output format. Then have a fluent speaker review it. Machine output is often grammatically fine but too long for the time available, and it can silently merge cues.

A note on Whisper: its built-in translate task only translates into English. It is useful for getting an English subtitle from foreign-language audio, but it will not give you French or Spanish subtitles from an English video, and it produces fresh timings rather than reusing yours. For how the model handles tasks, see how Whisper works.

Step 4: Re-check reading speed and line length

Once the text is back in the file, measure every cue again. The source-language numbers no longer apply. Our caption speed checker will flag cues with too many characters per second, and the caption reading speed guide explains the thresholds.

When a translated cue is too fast, you have three options, in this order of preference:

  1. Shorten the translation. This is almost always the right fix and does not affect timing.
  2. Extend the cue's out-time into a gap before the next cue, if there is one. Leave a small gap between cues so they do not appear to blink.
  3. Borrow time from a neighbouring cue by moving the boundary between them. Only do this when both cues still line up reasonably with the speech.

Avoid shifting a cue's in-time much earlier than the speech starts; text that appears before anyone speaks is confusing, and it spoils jokes and reveals.

Step 5: Check the technical details

Check Why it matters
File saved as UTF-8 Accented letters and non-Latin scripts turn into garbage otherwise
Timecode format unchanged SRT uses a comma before milliseconds (00:01:02,500), WebVTT uses a full stop (00:01:02.500)
Blank line between every cue A missing blank line can merge two cues into one
Two lines maximum per cue Longer blocks cover the picture and are hard to read
Right-to-left text displays correctly Arabic and Hebrew punctuation can appear at the wrong end of a line in some players
Font supports the script Matters for burned-in subtitles; missing glyphs show as boxes

If you need a WebVTT version of a translated SRT, or need to shift a whole file by a fixed offset, the SRT shifter tool handles both in the browser, without uploading the file anywhere.

Step 6: Name and upload each language separately

Name files with a language code so nothing gets mixed up: video.en.srt, video.de.srt, video.es.srt. On YouTube, each language is uploaded as its own subtitle track in YouTube Studio under Subtitles, where you choose the language before uploading the file. Viewers then pick it from the player settings. The exact menu layout changes from time to time, so check YouTube Help's "Add your own subtitles and captions" page if the steps look different.

If you are producing burned-in subtitles instead, you will need a separate video export per language. That trade-off is covered in open vs closed captions.

Quick checklist

  • Source subtitles corrected, synced and locked before translation starts
  • Translators receive text with cue IDs, never raw timecodes to edit
  • One source cue maps to exactly one translated cue
  • A glossary of names and terms goes to every translator
  • Reading speed re-checked in the target language
  • Long cues shortened first, re-timed only if necessary
  • Files saved as UTF-8 and named with a language code
  • A fluent speaker watches the video with the subtitles on before publishing

FAQ

Can I just run my SRT through a translation website?

You can, but check the output carefully. Some tools alter timecode punctuation, merge cues, or translate the cue numbers themselves. Translating the text separately and rebuilding the file, as shown above, avoids all three problems.

Why do my German subtitles feel rushed when the English ones were fine?

German text is often longer than the English equivalent, so the same cue duration has to hold more characters. Shorten the translation where you can, then extend cues into any available gaps.

Should the translated subtitles have the same number of cues as the original?

Yes, in almost every case. Keeping a one-to-one mapping means the timing you already checked still applies, and it makes review much easier. Only split or merge cues when you are prepared to re-time them.

Can Whisper translate English audio into other languages?

No. Whisper's translate task outputs English only. To create subtitles in other languages, transcribe in the original language, then translate the text as described in this guide.

Software, platform rules and settings change. We review our guides regularly, but always check the official documentation for the tools you use. Found an error? Email soubickdas@gmail.com. See our editorial policy.
A

AI Point Editorial

We build caption, transcription and video-workflow tools and write about what we learn doing it — practical, tested and free of hype.