# Cast voices in AI video — how a character speaks in Seedance and Hailuo renders

> How a character's voice rides a video render. Lines written inside a scene, the voice descriptor every engine with sound gets, the voice sample engines with audio references use, which engines do what, and how to write lines that fit a clip.

Published 2026-09-24 · eroq.ai — canonical: https://eroq.ai/blog/cast-voices-in-ai-video-seedance-hailuo


Video engines with a soundtrack invent a voice for anyone who speaks — a different one per clip, which is fine for a single shot and fatal for a film. eroq's answer is the character: give a cast member a voice in the studio, write their line in a scene, and on engines that take an audio reference the line comes out in *that* voice, scene after scene. This guide is what actually travels with a render, on which engines, and how to write dialogue that fits five seconds.

## What a character carries

A character in the studio has a sheet (the look), a persona (how they think and talk in chat) and a **voice** — a roster voice or one of your [clones](/blog/ai-voice-cloning-guide), assigned on the character's Voice tab. Three of those matter for video, and one does not:

- The **sheet** rides as a reference picture, so the face matches.
- The **voice descriptor** — a one-line description of the voice, « female voice, warm and low » — rides the prompt on every engine with sound.
- The **voice sample** — a short spoken sample of the assigned voice, rendered once and kept — rides as an audio reference on engines that accept one.
- The **persona** never rides a video prompt. The engine renders a picture; it does not need to know how she argues.

## Which engines do what

Three tiers, from the engine table:

- **Audio reference** — [Seedance 2.5](/models/seedance-2-5), [Seedance 2.0](/models/seedance-2-0), [Seedance 2.0 Fast](/models/seedance-2-0-fast), [Seedance 2.0 Mini](/models/seedance-2-0-mini), [Hailuo 3](/models/hailuo-3). The sample goes with the render; the character's lines come out in their voice. One sample per render — the first speaker in the scene who has a voice — so one speaking character per scene is the reliable case. A continuous shot, which opens on the previous scene's last frame, carries no references and so no sample: its voices ride as descriptors.
- **Sound, no audio reference** — [Veo 3.1](/models/veo-3-1) and its Fast and Lite tiers, [Sora 2 Pro](/models/sora-2-pro), [Kling 3.0](/models/kling-3) and Pro, [Wan 3.0](/models/wan-3), [Motion One](/models/eroq-motion-one). The descriptor rides the prompt; the engine casts a voice that matches the words. Consistent within a clip, not across clips.
- **No sound** — [Hailuo 3 Max](/models/hailuo-3-max), [Runway Gen-4.5](/models/runway-gen-4-5), [Grok Imagine](/models/grok-imagine). Lines still ride as direction (a mouth moving, a manner); nothing is heard. Add the voice in the edit from the [Voice tool](/studio/voice).

The Look panel in Cinema says which tier the film's engine is in. Sound can be switched off on any engine with a soundtrack, and off is cheaper where the engine charges for it.

## Writing the line

In Cinema, a scene's lines live inside its card: **+ Line** under the action, then a speaker, a manner and the line. The composer writes each one into the scene's own paragraph, right after the action — `Mina (whispering) says: "It's still lit."` — and closes the prompt with a labelled *Voices* sentence that says how each voiced character sounds, so it reads the same way every time. Fold open **What the engine reads** on the card to see the paragraph. Some rules that come from watching a few hundred renders:

- **One line per five seconds.** A short sentence, spoken, is three to four seconds. Two sentences need ten.
- **Manner over adjectives.** « whispering », « to herself », « not looking up » land; « emotionally » does not.
- **Punctuation is direction.** A dash is a pause, an ellipsis a trail-off, a question mark lifts the end. The [punctuation guide](/blog/punctuation-driven-voice-direction) applies to video as much as to speech.
- **Name the speaker in the action too** when two people are in frame: « Mina turns to Rook » before Mina's line, or the engine gives the line to whoever is closest to camera.
- **No stage directions inside the line.** The line is what is said; the manner is how. Mixing them gets the direction spoken aloud.

## A scene, end to end

The scene card:

```text
INT. LIGHTHOUSE — DUSK
@Mina reaches the top of the spiral stairs, out of breath, slow push-in
as she sees the lamp.

  MINA (whispering)   It's still lit.
```

Rendered on Seedance 2.0 Mini, 5 seconds, sound on — 45 credits. What the engine receives: the scene paragraph — the slugline, the action with the push-in, then her line with its manner; Mina bound to reference image 1, her sheet; the camera move and the look; and a Voices line telling the engine to speak with the voice in the reference audio — plus her voice sample as that audio reference. The clip has her face and her voice. Render scene two with another line and it is the same voice, because the same sample went — as long as scene two is not a continuous shot. A scene that opens on the previous frame sends no sample, so its voice comes from the descriptor; for a dialogue-heavy film where the voice matters more than the flow, turn **Continuous shots** off in the Look panel.

On Kling 3.0 the same scene costs 170 credits and the voice is Kling's guess from the descriptor — a good one, and a different one next scene.

## When the voice must be exact

For narration, a voice-over or a line that has to be *precisely* the clone, do not rely on the engine at all: render the line in the [Voice tool](/studio/voice) with the character's voice (3 credits per 100 characters), and drop the MP3 under the picture in Cinema's edit — with the clip's own sound muted or lowered. The engine's soundtrack is then ambience; the words are yours. This is also the only route on engines without sound.

## Limits, plainly

- One voice sample per render. A two-hander gets the first speaker's sample; the other gets the descriptor.
- The sample is a *reference*, not a script: the engine speaks in that voice; it does not lip-sync to a recording.
- Language follows the line. A French line in a voice cloned from English speech comes out accented, as it would in the Voice tool.
- Privacy: the sample is a signed link, valid long enough for the engine's queue, from your own storage.

## FAQ

### Does the character's persona affect how the line is delivered?

No. The persona is for chat. Delivery comes from the manner, the punctuation and the voice itself.

### Can two characters speak in one scene?

Both lines ride the prompt with their descriptors. Only one voice sample goes per render, so one of the two is the engine's cast. Split the exchange over two scenes for two exact voices, with Continuous shots off — a scene that opens on the previous frame sends no sample.

### Which engine keeps the voice most consistent?

The Seedance family and Hailuo 3, because they take the sample. Elsewhere the descriptor keeps the *kind* of voice consistent, not the voice.

### Is the voice sample charged?

No. It is rendered once per voice and kept; a render on an audio-reference engine sends it at no extra cost.

Assign a voice on a [character's Voice tab](/studio/characters), then write the line inside a scene card in [Cinema](/studio/cinema).
