Video
Text-to-video
Text-to-video is the generation of a video clip from a written prompt alone. The model synthesizes subject, motion, camera and lighting from the description, typically producing clips of a few seconds.
Modern text-to-video engines render coherent motion over 3 to 30 seconds. The prompt describes what is in frame, what moves, how the camera behaves and what the light does; the engine invents everything else. Quality varies by engine: some excel at human motion, others at physics or dramatic light, and only a few render sound.
Longer pieces are assembled from several clips — a storyboard — rather than generated in one pass.
On eroq
The video studio serves seven engines under one composer: Seedance 2.5 and 1.0, Kling 2.5 Turbo, Hailuo 02, Veo 3 Fast and the uncensored Motion One. Camera moves and a director's rack fold into the prompt server-side; Film mode chains scenes into one cut.
Questions
How long can text-to-video clips be?
From 3 to 30 seconds depending on the engine — Seedance 2.5 goes to 30; most others top out at 10–12.
What makes a good text-to-video prompt?
One paragraph naming subject, motion, camera, light and mood, 30–80 words, one action per clip.
More video terms