Short-form captions work best as one short line of two to four words, placed in the middle of the vertical frame, timed tightly to the speech and styled boldly enough to read on a small phone screen. They are a different craft from traditional subtitles: the goal is still that every word can be read, but the text is also part of the visual design. Below is how we plan, time and burn in captions for Shorts, Reels and TikTok.
Why short-form captions look different
Traditional subtitles sit at the bottom of a landscape frame, use two lines of up to around 40 characters and try to stay out of the way. Vertical short-form video breaks every one of those assumptions:
- The bottom of the frame is covered by the app's own interface: the caption and username overlay, music credit and, on the right, the like, comment and share buttons.
- The screen is narrow. A 1080-pixel-wide frame fits far fewer characters at a legible size than a 1920-pixel one.
- Many people watch muted, at least for the first few seconds, so the text carries the hook.
- Viewers are scrolling. Short chunks are easier to take in at a glance than full sentences.
So instead of two long lines, you show a phrase at a time.
The core rules
| Decision | Recommendation | Why |
|---|---|---|
| Words per caption | 2 to 4, split at natural phrase boundaries | Readable in one glance |
| Lines | One (two only for very short words) | Keeps the text block compact |
| Position | Roughly the middle third of the frame, horizontally centred | Clears the platform UI at top and bottom |
| Font | Heavy, clean sans-serif | Survives compression and small screens |
| Size | Large: roughly 60 to 90 px on a 1080 × 1920 frame, depending on the font | Legible at arm's length |
| Contrast | White or yellow text with a thick dark outline or shadow | Readable over any background |
| Minimum time on screen | Long enough to read; avoid flashes under about a third of a second | Very short flashes are missed |
The specific numbers are starting points rather than rules. Test on an actual phone, not just your editing monitor.
Where to put captions: safe zones
Each platform overlays its interface slightly differently, and those layouts change with app updates, so we do not quote exact pixel margins. At the time of writing, the consistent pattern across Shorts, Reels and TikTok is:
- The bottom fifth or so of the frame is often covered by the video description, username and audio details.
- A strip down the right-hand side holds the action buttons.
- The very top can hold the app's navigation or search elements.
Placing captions in the centre, or just below centre, avoids all three. If the speaker's face is in the middle of the frame, move the captions to sit just under the chin rather than across the mouth. The quickest check is to upload a private or draft test, then view it in each app and see what is covered. Our guide to video aspect ratios covers frame sizes if you are also cutting landscape versions.
Chunking: deciding where each caption breaks
Good chunking is what separates captions that feel punchy from ones that feel random. Split at the boundaries of meaning, not after an arbitrary word count.
Poor chunking:
SO THE BIGGEST
MISTAKE PEOPLE MAKE IS
THAT THEY
Better chunking:
THE BIGGEST MISTAKE
PEOPLE MAKE
IS THIS
Useful rules of thumb:
- Keep articles and prepositions with the word they belong to: "in the kitchen", not "in the / kitchen".
- Let a strong word stand alone for emphasis: "NEVER."
- End a chunk at punctuation in the script wherever possible.
- Do not create a chunk that briefly says something misleading on its own.
It is fine to drop filler such as "um", "like" and repeated false starts from captions on short-form content, but do not change what the speaker actually claims. The captions still need to be an honest record of the audio.
Timing with word-level timestamps
Short-form captions need precise timing because each chunk is on screen for well under a second in fast speech. Sentence-level timings from standard transcription are not enough; you need a start and end time for every word. Our guide to word-level timestamps explains how these are produced.
With the open-source Whisper command-line tool, recent versions can produce word timings and group them directly:
whisper clip.wav --model small --language en --word_timestamps True --max_words_per_line 3 --output_format srt
The grouping is purely by word count, so you will still want to adjust breaks by hand. A more controllable approach is to export word timings as JSON and group them yourself. This Python snippet reads Whisper's JSON output and starts a new chunk at punctuation, after three words or at a pause:
import json
MAX_WORDS, MAX_GAP = 3, 0.35
words = [w for seg in json.load(open("clip.json"))["segments"] for w in seg["words"]]
chunks, cur = [], []
for i, w in enumerate(words):
if cur and (w["start"] - cur[-1]["end"] > MAX_GAP):
chunks.append(cur); cur = []
cur.append(w)
if len(cur) >= MAX_WORDS or w["word"].strip()[-1:] in ".,!?":
chunks.append(cur); cur = []
if cur:
chunks.append(cur)
def ts(t):
h, rem = divmod(t, 3600); m, s = divmod(rem, 60)
return f"{int(h):02}:{int(m):02}:{s:06.3f}".replace(".", ",")
with open("clip_short.srt", "w", encoding="utf-8") as f:
for n, c in enumerate(chunks, 1):
text = " ".join(w["word"].strip() for w in c).upper()
f.write(f"{n}\n{ts(c[0]['start'])} --> {ts(c[-1]['end'])}\n{text}\n\n")
Always proofread the result. Recognition errors that are easy to overlook in a long subtitle file are glaring when a single wrong word fills the centre of the screen.
Styling with ASS and burning in with FFmpeg
Most short-form captions are burned into the video (open captions), because the style is part of the edit. ASS is the most practical format for this because it controls font, size, outline and position precisely; see SRT vs WebVTT vs ASS for background.
A style block for a 1080 × 1920 video with bold white text, a thick black outline and captions sitting a little below centre:
[Script Info]
ScriptType: v4.00+
PlayResX: 1080
PlayResY: 1920
[V4+ Styles]
Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding
Style: Short,Montserrat Black,78,&H00FFFFFF,&H0000FFFF,&H00000000,&H80000000,-1,0,0,0,100,100,0,0,1,6,2,2,90,90,720,1
Alignment 2 is bottom centre, and MarginV: 720 lifts the text 720 pixels up from the bottom edge, placing it in the lower-middle of the frame, clear of the bottom overlay. Left and right margins of 90 keep long chunks away from the edges.
You can convert your chunked SRT to ASS, replace its style section with the one above, then burn it in:
ffmpeg -i clip_short.srt clip_short.ass
ffmpeg -i clip.mp4 -vf "ass=clip_short.ass" -c:v libx264 -crf 18 -preset slow -c:a copy clip_captioned.mp4
The font must be installed on the machine running FFmpeg, or libass will substitute another. Only use fonts whose licence allows use in commercial video.
Word highlighting
The popular "active word" effect highlights each word as it is spoken. ASS supports this with karaoke tags: {\k40} holds the next syllable or word for 40 centiseconds, switching it from the secondary colour to the primary colour. A chunk using the style above, where words start yellow (secondary) and turn white as they are spoken, looks like this:
Dialogue: 0,0:00:03.20,0:00:04.10,Short,,0,0,0,,{\k30}THE {\k25}BIGGEST {\k35}MISTAKE
The karaoke durations come straight from the word timings. Use the effect with restraint: highlighting that is out of sync with the voice is more distracting than no highlighting at all.
Common mistakes
- Captions hidden under the app interface. Always check on the actual platform.
- Text too small or too thin for a phone screen, particularly with light outlines.
- Too many words per chunk, so viewers never finish reading before it changes.
- Emoji and colour on every line. Emphasis only works if it is occasional.
- Uncorrected recognition errors, especially names, brands and numbers.
- Captions flashing faster than people can read. Short-form chunks still need reading time; our caption reading speed guide explains the principles.
Quick checklist
- Transcribe with word-level timestamps and proofread every word.
- Chunk into 2 to 4 words at natural phrase boundaries.
- One line, centred, in the middle of the frame and clear of the platform UI.
- Heavy sans-serif, large size, strong outline or shadow.
- Highlight words only if the timing is accurate.
- Test on a phone in each app before publishing.
- Where the platform allows it, also add a caption file or turn on the platform's captions, so viewers who need captions can get them in their preferred style.
FAQ
Should I burn in captions or use the platform's auto-captions?
For short-form, burned-in captions give you control over style and accuracy, and they show even when a viewer has captions switched off. Where possible, also provide switchable captions, because they can be resized and read by assistive technology. Open vs closed captions covers the trade-offs.
Is all-caps text harder to read?
Long passages in capitals are generally harder to read, but for two to four words at a large size the difference is small, and capitals help with visual impact. Sentence case is a reasonable choice for calmer or educational content.
How many words per second is too fast for short-form captions?
There is no fixed figure, but if a chunk is on screen for less than about a third of a second, many viewers will miss it. Merge very short chunks with their neighbours instead.
Can I use the same captions for YouTube Shorts, Reels and TikTok?
Usually yes, provided the captions sit in the centre area that none of the apps cover. Check each platform, because their interfaces differ and change over time.