blog/tutorials·Sep 2, 2026·6 min·by the eroq team

Multi-language voice-over guide — 13 languages, one script

How to narrate a film in 13 languages with AI voice — script prep, punctuation that drives delivery, casting a voice once, and the credit math per character.


Localizing a short film used to mean thirteen bookings, thirteen studios and thirteen invoices. It now means thirteen text fields, which sounds easier than it is — the recording was never the hard part. The hard part is that a script written for one language, read at one cadence, laid under a picture cut to that cadence, does not survive translation unless you plan for it. This guide is the prep work, the punctuation, and the arithmetic.

What the voice studio actually gives you

The voice studio is text-to-speech with a curated roster — Aria, warm and intimate; Orion, low and calm — reading in thirteen languages, plus voice cloning from your own audio, speed and expressiveness controls, and delivery driven by the punctuation of the text you paste. Output is MP3.

Two models, same roster, same clones:

  • Voice One — the expressive flagship, 3 credits per 100 characters.
  • Voice Turbo — faster, 2 credits per 100 characters.

Because they share voices, you can draft every language on Turbo and re-read the keepers on Voice One without re-casting anything. That is a third off the exploratory phase for free.

Prep the script before you cast anything

Four things to fix while the script is still text.

Write for the ear. Sentences a reader forgives, a listener loses. Aim for one idea per sentence and put the verb early. If you cannot say a line in one breath, the engine will not fake one.

Spell out everything a machine has to guess. Numbers, units, currencies, dates, acronyms, product SKUs. "1080p" and "2026" and "approx." are typography, not speech; write what you want heard. This is the single largest source of wrong-sounding takes in any language.

Budget for expansion. The same sentence runs a different length in every language — sometimes noticeably longer than your English original. If your picture is cut to the English read, the other twelve will be fighting a wall. Leave air at the end of each beat, or cut the picture to the longest read and let the short ones breathe.

Translate the meaning, not the cadence. A word-for-word translation of a line written for English rhythm reads like a hostage note in most other languages. Have each language's script written as if it were the original, then check it fits the beat.

Punctuation is the direction

There is no performance dial hiding in the prompt. Delivery comes from punctuation, sentence length and the speed and expressiveness controls, and that is genuinely most of what you need:

  • Period. A full stop and a reset of the pitch.
  • Comma. A short breath. Two commas around a clause is the cheapest way to slow a line down.
  • Ellipsis. A trailing, unfinished thought. Powerful and easy to overuse — one per paragraph, maximum.
  • Em dash. An interruption, harder than a comma.
  • Question mark. A lift at the end of the line. Rhetorical questions are the easiest way to get intonation into flat narration.
  • Paragraph break. A real pause. Use blank lines to structure a read the way you would space a poem.

Per language, the rule is the same and the marks are not. Write each script with its own language's punctuation, typed the way a native typographer types it, and let those marks do the steering. Do not transplant English comma habits into a language that breathes somewhere else, and do not strip a language's own marks to make the file look tidy — you are deleting the direction. The language picker in the studio is the authority on which languages a given voice reads.

A complete narration block, punctuated for delivery rather than for grammar class:

Nobody comes to the twelfth floor after seven. The lifts still run — they just stop opening. She learned that on her third night, alone, with a security badge that worked on every door except one. And the one that did not open? It was not locked. It was held.

Six sentences, one em dash, one ellipsis avoided on purpose, one rhetorical question, ninety seconds of picture. Paste it, hit generate, and adjust speed before you touch anything else.

Cast once, reuse forever

Pick the voice in one sitting, for all thirteen languages, and then stop revisiting the decision. A roster voice is the simplest path. A clone made from your own audio is the differentiated one — voice cloning takes your recording once and the clone is then available like any roster voice, including on both speech models.

If you work with a recurring cast, attach the voice to the character record rather than keeping it in your head. A character carries a voice, so the same person sounds like the same person across a film, a trailer and a social cutdown. There is more on picking and directing voices for a cast in TTS for AI characters.

The arithmetic, per character and per language

Billing is per 100 characters of text, which makes budgeting unusually honest.

A 1,200-character narration — about ninety seconds of read — costs 36 credits on Voice One and 24 on Voice Turbo. Across thirteen languages that is 468 credits on Voice One, 312 on Turbo. At entry-pack rates, where a credit is roughly a cent, the whole localized narration lands under five dollars.

Now a scene with a speaking cast. Three characters, 300 characters of dialogue each, is 900 characters — 27 credits per language on Voice One, or 9 credits per character per language. Thirteen languages of that scene: 351 credits. A Hobby plan's 1,800 monthly credits therefore covers three full thirteen-language passes of a ninety-second narration with room for a fourth, and monthly credits roll over.

The expensive mistake is not the model, it is the re-read. Every time you revise a line after localizing, you pay for thirteen re-renders. Lock the script, then localize.

Check a language you do not speak

You will ship a take in a language you cannot hear mistakes in. Two cheap safeguards:

Transcribe your own output. Scribe One is speech-to-text at 5 credits per request for audio up to 8 MB, with automatic language detection. Run your generated MP3 back through it and read what the engine actually said. Mangled names, swallowed numbers and a mis-detected language all show up immediately in text.

Keep a pronunciation glossary. Proper nouns, brand names and invented place names are where every language goes wrong. Write the phonetic spelling you want into the script itself — the engine reads what you type, so type what you want heard.

Then composite

eroq does not trim, mix or grade, and there is no lip-sync — narration goes over picture, not into a mouth. Export the MP3, line it up in Premiere or After Effects against the cut, and duplicate the sequence per language. Every generation is saved automatically in your Library with its recipe, so a re-export six weeks later does not mean rewriting the settings. If you are wiring this into a pipeline instead of a timeline, the same voices are one POST away in the speech docs.

FAQ

How many languages can the voices read?

Thirteen. Both roster voices and your own clones read them, and the picker in the voice studio lists exactly which ones.

Do I need to clone a voice for every language?

No. Clone once and the clone is reusable — across languages, across projects and on both speech models.

What does a localized narration actually cost?

Voice One is 3 credits per 100 characters, Voice Turbo is 2. A 1,200-character script in thirteen languages is 468 credits on Voice One, 312 on Turbo.

Script locked? Open the voice studio and cast it once.

Tagsvoice-overtext-to-speechlanguagesvoice-cloningtutorial

Make this with the models behind the post — start with 50 free credits , or browse every engine and its price .