Video

Lip-sync

Lip-sync animates a face so the mouth matches an audio track you supply. It is what turns a still portrait or a silent clip into someone appearing to speak your words.

It is a distinct capability from video generation: the model is conditioned on the audio waveform and drives mouth shapes frame by frame. Tools that offer it generally take a face plus a voice file and return a new clip.

Generating speech inside the shot is a different thing entirely — there the engine invents both the performance and the sound, and you cannot choose the voice.

On eroq

eroq has no lip-sync tool. Veo 3 Fast can render a spoken line as part of the clip it generates, but you cannot lay a separate voice track over an existing clip and have the mouth match. Narration from the voice studio is laid under the picture in your own editor.

Questions

Can eroq make a character speak my script?

Not with matching mouth movement. You can render the voice from your script in the voice studio and the picture separately, then combine them in an editor — the face will not be synced.

What does Veo 3 Fast actually do with dialogue?

It generates the line as part of the shot, voice included, from your prompt. You do not choose the voice and you cannot supply the audio, so it suits incidental speech rather than scripted narration.

Is lip-sync on the roadmap?

We do not publish roadmap promises. What is true today is that no endpoint and no studio control does it.