AI Video and Captioning Glossary: 54 Terms Explained from A to Z

Plain-English A–Z glossary of AI video, captioning and transcription terms, from aspect ratio and bitrate to Whisper, word error rate and seeds.

Creator WorkflowBy AI Point EditorialUpdated 24 September 20269 min read

This glossary explains the terms you will meet when you make AI-assisted video, captions and transcripts, in plain English and in alphabetical order. Each definition is short and practical: what the word means, and why it matters when you are actually editing, exporting or captioning something.

We have kept to terms that come up again and again in real production work. Where a term deserves a proper explanation, we link to the full guide.

How to use this glossary

  • Skim the headings to find a letter, then read the definition in the list.
  • Related terms are grouped where it helps: for example, codec, container and bitrate all appear, and each definition tells you how it relates to the others.
  • Platform rules change. Where a definition touches a platform policy, check the platform's own help pages before relying on it.

A–C

  • ASR (automatic speech recognition): Software that turns spoken audio into text. Also called speech-to-text. Whisper is one well-known ASR model.
  • Aspect ratio: The shape of the frame, written as width to height. 16:9 is standard landscape YouTube; 9:16 is vertical for Shorts, Reels and TikTok. See our guide to video aspect ratios.
  • ASS/SSA: Advanced SubStation Alpha, a subtitle format that supports fonts, colours, positioning and karaoke-style effects. Popular for stylised burned-in captions.
  • Audio description: A narrated track describing important visual information for blind and partially sighted viewers. Different from captions, which cover the audio.
  • Bitrate: How much data per second a video or audio stream uses, usually in megabits per second (Mbps) for video and kilobits per second (kbps) for audio. Higher bitrate generally means better quality and bigger files, up to a point.
  • Burned-in captions: Captions drawn permanently into the picture, so they cannot be switched off. Also called open captions. Compare closed captions in open vs closed captions.
  • Caption: Text that represents speech and relevant sounds (such as "[door slams]"), intended for viewers who cannot hear the audio. Often used interchangeably with subtitles.
  • CFR / VFR: Constant frame rate and variable frame rate. Phone footage is often VFR, which can cause audio drift and caption sync problems in some editors. Converting to CFR before editing avoids this.
  • Character consistency: Keeping the same person, creature or object looking identical across several AI-generated images or clips. Usually achieved with reference images, fixed descriptions and, sometimes, LoRAs.
  • Characters per second (CPS): A reading-speed measure for captions: the number of characters in a caption divided by how long it stays on screen. Too high and viewers cannot keep up.
  • Closed captions (CC): Captions delivered as a separate track that viewers can turn on or off, such as an SRT or VTT file uploaded to YouTube.
  • Codec: The method used to compress and decompress video or audio. Common video codecs include H.264 (AVC), H.265 (HEVC), VP9 and AV1; common audio codecs include AAC and Opus.
  • Container: The file wrapper that holds video, audio, subtitle and metadata streams together, such as MP4, MOV or MKV. The container is not the same as the codec: an MP4 file might contain H.264 or H.265 video.
  • Content ID: YouTube's automated system that matches uploaded videos against reference files supplied by rights holders. A match can lead to a claim, blocked video or shared revenue.
  • CRF (constant rate factor): A quality-based encoding setting in encoders such as x264 and x265. Lower numbers mean higher quality and larger files; for x264 the default is 23.

D–H

  • Denoising: In AI image and video models, the step-by-step process of turning random noise into a picture. In editing, it also means removing visual grain or audio hiss.
  • Diffusion model: A type of generative model that learns to create images or video by reversing a noise-adding process. Most current text-to-image tools are diffusion-based.
  • Diarisation (speaker diarisation): Working out who spoke when in a recording, so a transcript can be labelled "Speaker 1", "Speaker 2" and so on.
  • Drop-frame timecode: A timecode method used with 29.97 fps video that skips certain frame numbers so the timecode stays in line with real clock time. Not needed at 25 fps.
  • Forced alignment: Matching a known, correct transcript to audio to get precise timings for each word or line. Explained in word-level timestamps.
  • Frame interpolation: Generating in-between frames to create smoother motion or slow motion. It can produce warping on fast or complex movement.
  • Frame rate: How many still frames are shown per second (fps). The UK's broadcast heritage is 25 and 50 fps; much online content is 24, 30 or 60 fps.
  • GOP (group of pictures): The structure of keyframes and dependent frames in compressed video. Long GOPs compress well but are harder for editing software to scrub through.
  • Hallucination: When an AI model produces confident output that is not supported by its input, such as a transcript containing words that were never spoken, often during silence or music.

I–L

  • Image-to-image (i2i): Generating a new image using an existing image as a starting point or guide, for example restyling a sketch or changing a character's pose.
  • Image-to-video: Animating a still image into a short clip using an AI video model.
  • Inpainting: Regenerating only a selected area of an image, such as fixing a hand or removing an object, whilst keeping the rest unchanged.
  • Keyframe: Two meanings. In animation, a point where you set a value (position, scale) that the software interpolates between. In compression, a full frame (I-frame) that other frames are predicted from.
  • Ken Burns effect: Slow panning and zooming across a still image to create a sense of movement. See the Ken Burns effect.
  • Latent space: The compressed internal representation a generative model works in. Many image models create a "latent" first and decode it into pixels at the end.
  • LoRA (low-rank adaptation): A small add-on file that fine-tunes a larger AI model for a specific style, character or object without retraining the whole model.
  • LUFS: Loudness Units relative to Full Scale, a measure of perceived loudness. UK and European broadcast work to EBU R 128, which targets −23 LUFS integrated. Streaming platforms apply their own normalisation, so check current guidance for each one.

M–R

  • Negative prompt: Text telling an image or video model what to avoid, such as "blurry" or "extra fingers". Not every model supports it.
  • Outpainting: Extending an image beyond its original borders, for example turning a square image into a 16:9 frame.
  • Prompt: The text instruction given to a generative model. Good prompts describe subject, setting, lighting, framing and style clearly.
  • Proxy: A lightweight, lower-resolution copy of high-resolution footage used for smooth editing, swapped back for the original at export.
  • Reference image: An image supplied alongside a prompt so the model keeps a face, outfit, product or style consistent.
  • Resolution: The pixel dimensions of a frame, such as 1920×1080 (1080p) or 3840×2160 (4K UHD).

S–T

  • SDH (subtitles for the deaf and hard of hearing): Subtitles that include speaker identification and non-speech sounds, not just dialogue.
  • Seed: The number that sets the starting noise for a generation. Reusing the same seed, prompt and settings on the same model usually reproduces a very similar result.
  • Speech-to-text: See ASR.
  • SRT: SubRip Text, the most widely supported caption format: numbered cues, a start and end timestamp, then the text. See what is an SRT file.
  • Storyboard: A sequence of sketches or frames planning each shot before production. With AI video, a storyboard keeps shots consistent and saves wasted generations.
  • Synthetic content disclosure: A platform label telling viewers that realistic content has been altered or generated. YouTube asks creators to disclose certain realistic synthetic content at upload.
  • Text-to-image / text-to-video: Generating an image or clip from a written prompt alone.
  • Timecode: A time reference in hours, minutes, seconds and frames (HH:MM:SS:FF) used to identify exact frames.
  • Transcode: Converting a file from one codec, bitrate or format to another. Every lossy transcode loses a little quality.
  • TTS (text-to-speech): Software that reads text aloud in a synthetic voice, often used for faceless channel narration.

U–Z

  • Upscaling: Increasing resolution. Traditional upscaling interpolates pixels; AI upscaling invents plausible detail, which can look sharp but occasionally wrong.
  • Voice cloning: Creating a synthetic voice modelled on a real person's recordings. Only clone a voice with that person's clear permission.
  • VTT (WebVTT): Web Video Text Tracks, the caption format used by HTML5 video. Similar to SRT but starts with a WEBVTT header, uses a full stop before milliseconds and supports positioning.
  • Whisper: An open-source speech recognition model released by OpenAI, available in several sizes that trade speed for accuracy.
  • Word error rate (WER): The standard ASR accuracy measure: substitutions plus deletions plus insertions, divided by the number of words in the correct transcript. Lower is better.
  • Word-level timestamps: Start and end times for each individual word, used for karaoke-style and short-form captions.

Terms that are easy to mix up

Often confused The difference
Codec vs container The codec compresses the streams; the container holds them. MP4 is a container, H.264 is a codec.
Captions vs subtitles Captions include relevant sounds for viewers who cannot hear; subtitles traditionally assume the viewer can hear and just needs the words or a translation.
Burned-in vs closed Burned-in text is part of the picture; closed captions are a separate, switchable track.
Keyframe (animation) vs keyframe (encoding) One is a value you set in your editor; the other is a full frame inside the compressed file.
Bitrate vs resolution Resolution is how many pixels; bitrate is how much data describes them each second.

A quick way to check the codec, container, frame rate and bitrate of any file is FFmpeg's ffprobe:

ffprobe -v error -show_entries format=format_name,bit_rate:stream=codec_name,width,height,r_frame_rate -of default=nw=1 input.mp4

Key takeaways

  • Learn the difference between codec, container and bitrate first; most export problems come from mixing them up.
  • For captions, the core vocabulary is SRT, VTT, CPS and burned-in vs closed.
  • For AI generation, prompt, seed, reference image and negative prompt control most of what you get.
  • For transcription, ASR, WER, diarisation and forced alignment describe what a tool does and how well.

FAQ

Is a caption the same as a subtitle?

In everyday use, often yes. Strictly, captions (and SDH) include sound effects and speaker labels for viewers who cannot hear the audio, whilst subtitles usually carry only dialogue or a translation.

What is the difference between H.264 and MP4?

H.264 is a video codec, the compression method. MP4 is a container file that can hold H.264 video alongside AAC audio and other streams.

Why does the same prompt give different AI images?

Because the seed changes. Most tools pick a random seed each time; fix the seed and keep the model and settings the same if you want repeatable results.

What is a good word error rate?

It depends on the audio. Clean, single-speaker studio narration can give very low WER, whilst noisy, overlapping speech will be far higher. Test on a sample of your own recordings rather than trusting a headline figure.

Software, platform rules and settings change. We review our guides regularly, but always check the official documentation for the tools you use. Found an error? Email soubickdas@gmail.com. See our editorial policy.
A

AI Point Editorial

We build caption, transcription and video-workflow tools and write about what we learn doing it — practical, tested and free of hype.