# Text-to-video

> Text-to-video is the generation of a video clip from a written prompt alone. The model synthesizes subject, motion, camera and lighting from the description, typically producing clips of a few seconds.

Modern text-to-video engines render coherent motion over 3 to 30 seconds. The prompt describes what is in frame, what moves, how the camera behaves and what the light does; the engine invents everything else. Quality varies by engine: some excel at human motion, others at physics or dramatic light, and only a few render sound.

Longer pieces are assembled from several clips — a storyboard — rather than generated in one pass.

## On eroq

The video studio serves seven engines under one composer: Seedance 2.5 and 1.0, Kling 2.5 Turbo, Hailuo 02, Veo 3 Fast and the uncensored Motion One. Camera moves and a director's rack fold into the prompt server-side; Film mode chains scenes into one cut.

## FAQ

### How long can text-to-video clips be?

From 3 to 30 seconds depending on the engine — Seedance 2.5 goes to 30; most others top out at 10–12.

### What makes a good text-to-video prompt?

One paragraph naming subject, motion, camera, light and mood, 30–80 words, one action per clip.

## Related

- https://eroq.ai/tools/text-to-video — Text to video
- https://eroq.ai/glossary/image-to-video — Image-to-video
- https://eroq.ai/glossary/storyboard — Storyboard

---

This page as HTML: https://eroq.ai/glossary/text-to-video · Studio: https://eroq.ai/studio · Docs: https://eroq.ai/docs · Machine index: https://eroq.ai/llms.txt
