blog/tutorials·Sep 10, 2026·6 min·by the eroq team
How to add an AI voice-over to your video, start to finish
Write the script, render it on an expressive voice, time it to the clip and lay it under the cut in Premiere with the eroq panel. A practical workflow.
A silent AI clip is a wallpaper. The same clip with thirty seconds of narration is a piece of content someone watches to the end. Voice-over is the cheapest upgrade available in the whole pipeline — a hundred-word script is a few credits, against a hundred or more for the footage it sits under.
Before the steps, one thing to be clear about, because it shapes everything else: this is narration, not lip-sync. Nothing here matches a mouth to your audio. You are building a separate voice track that plays over your footage, the way a documentary or a faceless channel does. Write for that and it works beautifully. Expect a talking mouth and you will be disappointed.
Step 1 — write the script where rewriting is cheap
Write in the text studio rather than in the voice box. Two reasons: you get a model that is good at prose, and you get a thread you can iterate in without re-rendering audio each time.
Pick the register deliberately. Scene produces immersive prose — the right choice for narration over a film. Messaging produces short, texting-style lines — the right choice for a character speaking, or for punchy ad copy. Completions cost 1 or 3 credits depending on the model, so drafting is effectively free next to the footage.
Two rules that matter more than the writing itself:
- Write shorter than the clip. Silence at the head and tail of a shot is a gift to your edit. A voice track that fills every frame has nowhere to breathe.
- Write in sentences you would say out loud. If you stumble reading it, the engine will too.
Step 2 — pick a voice
Open the voice studio. The curated roster includes Aria, warm and intimate, female, and Orion, low and calm, male, across 13 languages, alongside the rest of the catalog. Search by name, category or language, and listen through the roster before you commit to one.
If you need a specific voice that is not on the roster, clone one from your own audio — see the voice cloning tool page for what makes a usable sample. Clones then behave exactly like roster voices everywhere in the studio. A saved Character can also carry a voice, roster or cloned, so a recurring narrator stays the same across a series without you remembering which one you used.
Step 3 — render it
Two models, same roster, same clones:
- Voice One — the expressive flagship, 3 credits per 100 characters, MP3 out. Use it for anything a viewer will hear more than once.
- Voice Turbo — faster, 2 credits per 100 characters. Use it for drafts, timing tests, and long-form where the budget is the constraint.
Both take Speed and Expressiveness controls. Speed is the one to reach for when a line is three seconds too long for its shot. Expressiveness is worth testing in both directions — narration usually wants less than you think, character lines usually want more.
Every render lands in your library like everything else, so you can come back to the take you liked.
Step 4 — punctuation is the direction
There is no separate emotion parameter. Delivery is driven by the text itself, and specifically by punctuation, which means your script is the direction. The honest version of what that gets you:
- A period ends the thought and drops the pitch.
- A comma is a short lift, not a stop.
- An em dash — like this one — holds the beat longer than a comma.
- Ellipses trail off and soften the ending.
- A question mark lifts, even mid-paragraph.
- Paragraph breaks are the longest pause you can ask for.
A script written for delivery rather than for the page:
He left the keys on the table. No note — nothing.
I stood there for a long time, holding a coffee that had gone cold, trying to work out which part of this I had agreed to.
And the strangest thing? The door was still unlocked.
Read that aloud and you can hear where it breathes. That is the test. If you want a longer pause, use a paragraph break; if a line lands flat, the fix is usually a dash where you wrote a comma.
What punctuation cannot do is bolt an emotion onto a sentence that does not have one. Write the feeling into the words.
Step 5 — time it to the clip
Do this before you open an editor, not after.
- Read the script aloud with a stopwatch. That number is close enough to plan with.
- Choose the clip length to fit it — engines render from 3 up to 30 seconds depending on the model, and your plan caps the maximum at 10, 15, 20 or 30 seconds.
- Render the voice on Turbo first and listen against the clip.
- If it is long, cut words before you raise the speed. Speed above a light nudge starts to sound like speed.
- Re-render the final on the flagship voice once the words are settled.
For sequences, do it per shot instead of per film. One long narration track over six clips means every edit change forces a re-render; six short lines mean a change costs one. This is also how faceless YouTube workflows stay maintainable at volume.
Step 6 — lay it under the cut
Download the MP3 and drop it into whatever you edit in. If that is Premiere Pro or After Effects, the eroq panel saves you the round trip: generate video, images or speech from inside the panel and the renders import straight into the open project, with files saved under your Documents folder. Generate a line, drop it on the timeline, adjust, generate the next one — without leaving the edit.
Then mix like a human. Voice sits on top, everything else ducks under it, and the last word should land before the last frame. If a shot carries its own audio from an engine that renders sound natively, ride that bed under the narration rather than muting it — ambience under a voice is most of what "produced" sounds like.
The speech docs cover the same thing from the API side if you want to script it, and the video studio is where the footage half of this lives.
FAQ
Can eroq lip-sync a character to my voice-over?
No. Lip-sync is not a feature. The voice track is generated independently of the footage, which makes it ideal for narration, ads and faceless content — write and cut on that assumption.
How much does a voice-over cost?
3 credits per 100 characters on the flagship voice and 2 on the faster one, MP3 out. A 600-character script is roughly 18 or 12 credits, which is small next to the clip it plays over.
Can I use the same narrator across a whole series?
Yes. Pick a roster voice or clone one, and attach it to a saved Character so the same voice is selected every time. Clones work anywhere a roster voice does.
Write it, say it, cut it. Open the voice studio — new accounts start with 50 free credits.
Make this with the models behind the post — start with 50 free credits , or browse every engine and its price .