FFmpeg is a free, open-source command-line tool that can trim, join, convert, resize, caption and compress almost any video or audio file. You do not need to learn all of it: a dozen commands cover most day-to-day creator jobs. Below are the ones we use constantly when building caption and video tools, each with a short explanation of what every part does and the mistakes to avoid.
Before you start
Install FFmpeg from the official download page at >ffmpeg.org, which links to trusted builds for Windows, macOS and Linux. Check it works:
ffmpeg -version
A few conventions used throughout:
-isets an input file. Options before-iapply to that input; options after it apply to the output.-c copycopies streams without re-encoding. It is fast and lossless, but you cannot change the picture or sound while copying.-c:vand-c:aset the video and audio codec separately.-vfapplies a video filter;-afapplies an audio filter.- On Windows, the commands work the same in PowerShell or Command Prompt, but use double quotes around paths with spaces.
Always work on copies, and never overwrite your only original. Keep raw files in a separate folder that nothing writes into.
1. Inspect a file first
Before changing anything, find out what you have.
ffprobe -v error -show_entries stream=codec_type,codec_name,width,height,r_frame_rate,sample_rate -of default=nw=1 input.mp4
This prints the codec, resolution, frame rate and audio sample rate of each stream. Knowing the frame rate up front avoids a lot of sync problems later; see frame rates explained for why.
2. Trim a clip
Fast, no re-encode:
ffmpeg -ss 00:01:30 -i input.mp4 -t 45 -c copy clip.mp4
-ss before -i seeks to 1 minute 30 seconds, and -t 45 keeps 45 seconds. Because the streams are copied, the cut can only start on a keyframe, so the clip may begin slightly early or with a brief frozen frame.
Frame-accurate, with re-encode:
ffmpeg -ss 00:01:30 -i input.mp4 -t 45 -c:v libx264 -crf 18 -preset medium -c:a aac -b:a 192k clip.mp4
Use this when the exact start frame matters, such as a clip for a Short.
3. Join clips together
Create a text file called list.txt:
file 'part1.mp4'
file 'part2.mp4'
file 'part3.mp4'
Then:
ffmpeg -f concat -safe 0 -i list.txt -c copy joined.mp4
This is the concat demuxer. It only works reliably with -c copy when every clip has the same codec, resolution, frame rate and audio settings, which is usually true for clips from the same camera or the same export preset. If they differ, re-encode each clip to matching settings first, or replace -c copy with encoding options. -safe 0 allows full or unusual file paths in the list.
4. Extract the audio
Keep the original audio exactly (if it is AAC, as most MP4s are):
ffmpeg -i input.mp4 -vn -c:a copy audio.m4a
Convert to WAV for editing or transcription:
ffmpeg -i input.mp4 -vn -c:a pcm_s16le -ar 48000 -ac 2 audio.wav
-vn drops the video. For speech-to-text, a mono 16 kHz WAV (-ac 1 -ar 16000) is smaller and matches what many speech models use internally, though most tools resample for you.
5. Normalise loudness
ffmpeg -i input.mp4 -c:v copy -af loudnorm=I=-14:TP=-1.5:LRA=11 -ar 48000 -c:a aac -b:a 192k normalised.mp4
The loudnorm filter implements the EBU R128 loudness method. I is the integrated loudness target in LUFS, TP the true-peak ceiling in dBTP, and LRA the loudness range. Around -14 LUFS is a common target for online platforms; UK broadcast delivery uses -23 LUFS, so follow the specification if you have one.
Two practical notes:
loudnormworks internally at a high sample rate, so always set-ar 48000or you may get a 192 kHz output.- A single pass is fine for most creator work. For the most accurate result, run it twice: first with
print_format=jsonto measure, then feed the measured values back in. The >FFmpeg filter documentation explains the parameters.
6. Burn subtitles into the video
ffmpeg -i input.mp4 -vf "subtitles=captions.srt" -c:v libx264 -crf 18 -preset medium -c:a copy captioned.mp4
This permanently renders the captions onto the picture ("open" captions). To change the look:
ffmpeg -i input.mp4 -vf "subtitles=captions.srt:force_style='FontName=Arial,FontSize=22,Outline=2,MarginV=60'" -c:a copy captioned.mp4
Common mistakes:
- Windows paths. The filter's own syntax treats
:and\specially. The simplest fix is to run the command from the folder containing the subtitle file and use just the file name. Otherwise, use forward slashes and escape the drive colon, for examplesubtitles='C\:/Projects/captions.srt'. - No libass. Burning subtitles needs an FFmpeg build with libass. Most full builds include it.
- ASS files keep their own styling, which usually looks better than
force_styleon an SRT. See SRT vs VTT vs ASS.
Whether to burn captions in at all, rather than uploading a separate caption file, is a separate decision worth making deliberately.
7. Convert SRT to WebVTT
ffmpeg -i captions.srt captions.vtt
FFmpeg reads the SRT and writes a valid WebVTT file, including the WEBVTT header and dot-separated milliseconds. It does not keep SRT-specific quirks such as unusual formatting tags. For shifting timings as well as converting, the SRT shifter tool does both in the browser.
8. Reframe landscape to 9:16
Fit with black bars:
ffmpeg -i landscape.mp4 -vf "scale=1080:1920:force_original_aspect_ratio=decrease,pad=1080:1920:(ow-iw)/2:(oh-ih)/2:color=black,setsar=1" -c:v libx264 -crf 20 -c:a copy vertical.mp4
Crop to fill the screen (centred):
ffmpeg -i landscape.mp4 -vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1" -c:v libx264 -crf 20 -c:a copy vertical_crop.mp4
Blurred background behind the fitted video:
ffmpeg -i landscape.mp4 -filter_complex "[0:v]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,boxblur=20:1[bg];[0:v]scale=1080:-2[fg];[bg][fg]overlay=(W-w)/2:(H-h)/2,setsar=1" -c:v libx264 -crf 20 -c:a copy vertical_blur.mp4
scale=1080:-2 sets the width to 1080 and picks an even height that keeps the proportions. setsar=1 forces square pixels so nothing looks stretched. More on shapes in video aspect ratios.
9. Grab a thumbnail frame
ffmpeg -ss 00:00:12 -i input.mp4 -frames:v 1 -q:v 2 frame.jpg
-frames:v 1 saves a single frame; -q:v 2 is high JPEG quality (lower is better, 2–5 is sensible). For a lossless still to design over, use frame.png and drop -q:v. To let FFmpeg pick a representative frame from the start of the clip:
ffmpeg -i input.mp4 -vf thumbnail -frames:v 1 auto_thumb.png
10. Change the frame rate
ffmpeg -i input.mp4 -vf fps=25 -c:v libx264 -crf 18 -preset medium -c:a copy output_25p.mp4
The fps filter duplicates or drops frames to reach the target rate while keeping the duration and audio sync the same. It is also the standard fix for variable frame rate phone and screen recordings. It cannot create genuinely new motion, so converting 25 to 30 will show a slight stutter on movement.
11. Compress with CRF
ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset slow -c:a aac -b:a 128k -movflags +faststart compressed.mp4
CRF (constant rate factor) targets consistent visual quality rather than a fixed file size. For libx264, lower numbers mean higher quality and bigger files: 18 is close to visually lossless, 23 is the default, and 28 is noticeably softer. -preset slow compresses more efficiently than medium at the cost of encoding time. -movflags +faststart moves index data to the start of the file so web players can begin playback sooner.
For smaller files with H.265:
ffmpeg -i input.mp4 -c:v libx265 -crf 28 -preset medium -tag:v hvc1 -c:a aac -b:a 128k compressed_hevc.mp4
libx265's CRF scale is different, and 28 is its default. -tag:v hvc1 helps Apple devices recognise the file. For upload masters, keep quality high and let the platform compress; see YouTube export settings.
12. Turn an image sequence into a video
ffmpeg -framerate 25 -i frame_%04d.png -c:v libx264 -crf 18 -pix_fmt yuv420p sequence.mp4
%04d matches frame_0001.png, frame_0002.png and so on, which is why consistent numbering matters. -pix_fmt yuv420p ensures the result plays in browsers and phones. If your files are misnumbered, fix them first; see batch renaming image sequences or the Image Replicator.
Quick reference
| Job | Key options |
|---|---|
| Inspect | ffprobe -show_entries stream=... |
| Trim fast | -ss before -i, -t, -c copy |
| Join | -f concat -safe 0 -i list.txt -c copy |
| Extract audio | -vn -c:a copy or -c:a pcm_s16le |
| Loudness | -af loudnorm=I=-14:TP=-1.5:LRA=11 -ar 48000 |
| Burn captions | -vf subtitles=file.srt |
| SRT to VTT | ffmpeg -i in.srt out.vtt |
| 9:16 | scale=...:force_original_aspect_ratio=... plus pad or crop |
| Thumbnail | -frames:v 1 -q:v 2 |
| Frame rate | -vf fps=25 |
| Compress | -c:v libx264 -crf 23 -preset slow |
| Image sequence | -framerate 25 -i name_%04d.png -pix_fmt yuv420p |
FAQ
Why does my trimmed clip start with a frozen frame?
You used -c copy, which can only cut on keyframes. Re-encode for a frame-accurate cut, as in command 2.
Is re-encoding with FFmpeg lossy?
Yes, any encode with a lossy codec such as H.264 loses some detail. Keep CRF low (around 18) for intermediate files, and use -c copy whenever you are not changing the picture or sound.
Why does the concat command give errors or glitches?
The clips probably differ in codec, resolution, frame rate or audio format. Re-encode them to identical settings first, then join.
Can FFmpeg transcribe speech?
FFmpeg prepares the audio; a speech model such as Whisper does the transcription. FFmpeg can convert the audio to a format such a model accepts, as in command 4.