5 prompts · Voiceover narration
AI voiceover prompts and TTS script preparation
You want synthetic narration that does not sound synthetic.
TTS quality is mostly a script problem. Engines read punctuation as timing, mishandle numbers and acronyms, and flatten out on long sentences. Preparing the text properly removes most re-render cycles and is far cheaper than switching voices repeatedly.
Faceless documentary narrationAudiobook and course narrationMultilingual dubs of existing video
The prompts
1. TTS preparation pass
Claude or ChatGPT, then ElevenLabs or MurfConverting a script into engine-ready text
Prepare this script for text-to-speech narration.
1. Split any sentence over 22 words at a natural breath point.
2. Spell out numbers, dates, currencies, units, and acronyms as they should be spoken ("2026" → "twenty twenty-six", "API" → "A P I").
3. Give phonetic spellings in brackets for proper nouns and unusual words.
4. Mark pauses with ellipses where the meaning needs a beat.
5. Bold 5-8 words for emphasis across the whole script.
Do not change meaning. Output the prepared script plus a table of every substitution.
Script: [PASTE SCRIPT]In practice: This one pass eliminates the majority of re-renders.
2. Voice direction brief
Claude or ChatGPTChoosing and configuring the voice
My video is about [TOPIC], for [AUDIENCE], with a [TONE] feel.
Describe the ideal narrator: age range, accent, pace in words per minute, energy level, and warmth. Then give me the two or three settings I should adjust in a TTS tool (stability, similarity, speed, style exaggeration) and in which direction, and explain the trade-off each one makes.
In practice: Naming a target words-per-minute stops the default read from being too fast.
3. Emphasis and pacing markup
Claude or ChatGPTMaking a flat read dynamic
Mark up this narration for delivery. For each paragraph, note the emotional beat in one word, the intended pace (slow / steady / brisk), and which single word carries the sentence.
Then rewrite with pauses, emphasis, and sentence-length variation that produce that read. A short sentence after two long ones. Never three long sentences in a row.
Script: [PASTE SCRIPT]
In practice: Sentence-length variation does more for perceived naturalness than any voice setting.
4. Dub script for another language
Claude or ChatGPT, then Papercup or ElevenLabsLocalising narration to fit the original timing
Translate this narration into [LANGUAGE] for dubbing over existing video.
Critical: each segment must take approximately the same time to speak as the original. Where the natural translation runs long, tighten it rather than speeding up delivery. Keep proper nouns in their original form unless there is an established local version.
Output as a table: timestamp, original line, translated line, estimated spoken seconds.
Script: [PASTE SCRIPT WITH TIMESTAMPS]
In practice: Timing-matched translation is the whole job in dubbing — literal accuracy is secondary.
5. Diagnosing a bad render
Claude or ChatGPTFixing narration that sounds off
My TTS narration sounds [DESCRIBE — rushed, flat, robotic on certain words, wrong emphasis, unnatural pauses]. Here is the text I fed it: [PASTE].
Identify which parts of the text are causing that, specifically. Rewrite only those parts. Tell me which issues are text problems and which are voice-setting problems, because I can only fix the first kind here.
In practice: Separating text problems from settings problems saves an afternoon of slider-fiddling.
Tools that finish the job
Where these prompts go next
A prompt produces the script, the shot list or the markup. These are the tools that turn that output into a finished video.
Play.ht
AI voice generator with 900+ ultra-realistic voices for voiceovers, podcasts, and text-to-speech.
Creator from $31/mo; Pro from $49/mo
View tool profile →Papercup
AI dubbing and voice localization for video content.
Contact for pricing
View tool profile →