Writing for AI Voiceover: How to Script Narration That Text-to-Speech Reads Well

How to write narration scripts for AI and text-to-speech voiceover: pacing, punctuation, numbers, tricky UK names, pronunciation fixes and a checklist.

AI Video CreationBy AI Point EditorialUpdated 24 September 20268 min read

Writing for an AI voice is writing for the ear, with one extra constraint: the voice will read exactly what you type, including every ambiguity. Good TTS scripts use short sentences, spell out numbers and abbreviations the way you want them spoken, use punctuation to control pauses, and avoid words the engine is likely to mispronounce. Do that, and a modern voice sounds natural; skip it, and even the best voice stumbles.

Why text-to-speech needs its own style

A human narrator reads a script, notices "£2.5m" and says "two and a half million pounds". A text-to-speech engine has to guess. Most modern voices guess well most of the time, but "most of the time" is not good enough when you are rendering a 10-minute narration and do not want to re-listen to every line.

TTS engines also lack context. They cannot see your B-roll, do not know that "live" in this sentence is the adjective, and will not pause for effect unless the text tells them to. The script has to carry all of that information.

Pacing: plan your length before you write

A common rule of thumb for narration is around 150 words per minute, with calm documentary voices nearer 130 and energetic short-form voices nearer 170. The exact figure depends on the voice and its speed setting, so measure your own:

  1. Render a paragraph of about 150 words with your chosen voice and settings.
  2. Note the duration.
  3. Divide words by minutes to get your personal words-per-minute figure.

Use that number to budget scripts. If your voice reads at 150 wpm, a 60-second Short needs about 150 words, and an 8-minute explainer about 1,200. Writing to a word budget is far quicker than cutting audio afterwards.

Sentence structure

Keep sentences short and single-idea

Long sentences with several clauses are where TTS voices most often lose their intonation, rising where they should fall or running out of "breath". Aim for one idea per sentence, and vary length a little so the rhythm does not become a monotone.

Harder for TTS:

The bridge, which was designed in 1890 by an engineer who had never built anything longer than a railway viaduct, collapsed only nine years later, although nobody, at least officially, was blamed.

Easier for TTS:

The bridge was designed in 1890. Its engineer had never built anything longer than a railway viaduct. Nine years later, it collapsed. Officially, nobody was blamed.

The second version also gives your editor natural cut points for visuals.

Front-load the point

Listeners cannot re-read. Put the key information early in the sentence ("Nine years later, it collapsed") rather than at the end of a long build-up.

Write the way people speak

Contractions ("it's", "didn't", "you'll") usually sound more natural than the formal versions. Avoid written-only constructions such as "the former... the latter", bracketed asides, and "respectively".

Punctuation is your pause control

Most TTS engines treat punctuation as timing cues. The details vary by engine, but broadly:

Punctuation Typical effect
Comma Short pause, pitch often held
Full stop Longer pause, pitch falls
Question mark Rising or questioning intonation
Ellipsis (...) Hesitant or trailing pause (engine-dependent)
Dash (—) Short break, sometimes treated like a comma
New paragraph Often a slightly longer pause

Use them deliberately. If a line needs a dramatic beat before a reveal, a full stop works better than a comma. Avoid stacking exclamation marks; many voices either ignore them or over-act.

Some tools support SSML (Speech Synthesis Markup Language), a W3C standard that lets you insert exact pauses, emphasis and pronunciations with tags such as <break time="500ms"/>. Support varies a lot between providers, so check your tool's documentation. The full standard is published at >w3.org.

Numbers, dates, money and units

Write numbers the way you want them heard. This single habit fixes most TTS errors.

Written Risk Write instead
£2.5m "two point five m" two and a half million pounds
1,200 "one thousand two hundred" vs "twelve hundred" whichever you want, in words
03/04/2025 UK vs US date order the third of April, 2025
1990s usually fine the nineteen-nineties (if unsure)
10km "ten k m" ten kilometres
3x faster "three x faster" three times faster
24/7 "twenty-four slash seven" twenty-four seven
No. 10 "no ten" Number 10

Years are usually read correctly ("1066" as "ten sixty-six"), but test anything unusual. Phone numbers, codes and version numbers ("v2.1") should be written exactly as they should sound.

Pronunciation problems to watch for

Homographs

Words spelt the same but pronounced differently are a classic trap: "read" (present or past), "lead" (metal or guide), "live" (verb or adjective), "wind", "tear", "close", "bass", "minute". Engines usually infer correctly from context, but not always. If a line depends on one, listen to it, and rephrase if necessary ("the live broadcast" might become "the broadcast, shown live").

UK place and family names

Many voices, especially those trained mostly on American English, struggle with British names that are not spoken as spelt: Leicester, Worcester, Gloucester, Loughborough, Bicester, Keswick, Magdalen College, Featherstonehaugh. Options, in order of preference:

  1. Use a pronunciation dictionary or "lexicon" feature if your tool has one, so the fix applies everywhere.
  2. Use SSML phoneme tags, where supported.
  3. Respell phonetically in the script ("Lester" for Leicester). This works but can leak into captions if you generate them from the script, so keep a clean copy.

Acronyms and initialisms

Decide whether each one is spoken as a word ("NATO", "Ofcom") or letter by letter ("BBC", "NHS"). Engines usually handle famous ones; lesser-known ones may need spacing or dots ("N. H. S.") or spelling out in full on first use.

Foreign words and names

Test every one. If the voice cannot say it convincingly, consider a different phrasing or a brief on-screen label rather than repeated attempts.

Keep the script and captions in step

If you respell words for the voice, your captions must still show the correct spelling. We keep two versions: a clean script (used for captions and the description) and a voice script (with respellings and pause tweaks). Generating captions from the rendered audio with a speech-to-text model and then correcting against the clean script works well; our guide to word-level timestamps explains how to get precise timing for short-form captions.

It also helps to check caption speed once the voice is rendered. A fast voice can produce captions that are too quick to read comfortably; the caption speed checker and our guide to caption reading speed show what to aim for.

Ethics and disclosure

Only use voices you have the right to use. Stock voices in a commercial TTS tool come with licence terms; cloning a real person's voice needs their clear, informed consent. On YouTube, a realistic synthetic voice made to sound like a real, identifiable person falls under the platform's altered or synthetic content rules. See our guide to YouTube's AI disclosure for what needs labelling.

A workflow that works

  1. Draft the script to a word budget based on your measured words per minute.
  2. Read it aloud yourself once. Anything you stumble on, the voice probably will too.
  3. Convert numbers, dates, money and units to words.
  4. Mark or fix homographs, acronyms and difficult names.
  5. Render in paragraphs or scenes, not one giant block, so you can re-render a single line.
  6. Listen to every render at normal speed with the script in front of you.
  7. Fix errors in the script (or lexicon), not by patching audio.
  8. Generate captions and check them against the clean script.

For a wider view of how narration fits into a faceless production, see our faceless YouTube channel workflow.

Quick checklist

  • Word count matches your target length at your measured pace.
  • Sentences are short, one idea each, with the key point first.
  • All numbers, money, dates and units are written as they should be spoken.
  • Homographs, acronyms and UK names have been checked by ear.
  • Punctuation is used deliberately for pauses.
  • A clean script is kept separately for captions.
  • Voice licence is clear and any realistic likeness is disclosed.

FAQ

How many words is a one-minute voiceover?

Roughly 130–170 words, depending on the voice and speed setting. Measure your chosen voice once and use that figure for every script.

Should I use SSML?

If your tool supports it and you need exact pauses or pronunciations, yes. If it does not, punctuation, paragraph breaks and phonetic respelling get you most of the way.

Why does my AI voice sound flat in long paragraphs?

Long sentences and long unbroken blocks tend to flatten intonation. Split sentences, add paragraph breaks and render in shorter sections.

Can I just paste a blog post into a TTS tool?

You can, but written prose is not spoken prose. Expect to rewrite for the ear: shorter sentences, fewer asides, spoken-form numbers and a clear structure the listener can follow without seeing the text.

Software, platform rules and settings change. We review our guides regularly, but always check the official documentation for the tools you use. Found an error? Email soubickdas@gmail.com. See our editorial policy.
A

AI Point Editorial

We build caption, transcription and video-workflow tools and write about what we learn doing it — practical, tested and free of hype.