# AI video prompt guide — how to write prompts that actually render

> The subject, motion, camera, light and mood structure video engines follow, the mistakes that break them, before-and-after rewrites and the free Enhance button.

Published 2026-09-01 · eroq.ai — canonical: https://eroq.ai/blog/ai-video-prompt-guide


Most AI video prompts fail for structural reasons, not for lack of imagination. A video engine reads your text as a shot description; hand it a mood board, a plot summary or a pile of tags and it renders its best guess, which is usually a slow blur of the first noun. This guide is the structure we use in [the Video studio](/studio/video), the mistakes that break every engine we run, and how to fix a prompt in one pass.

## The five parts of a prompt that renders

A working video prompt is one flowing paragraph, 30 to 80 words, that answers five questions in roughly this order:

1. **Subject** — who or what, with two or three concrete visual details. "A woman in a yellow raincoat" beats "a mysterious figure".
2. **Motion** — what the subject does during the clip. One continuous action. Engines are good at "walks toward the camera" and bad at "walks, stops, turns, waves".
3. **Camera** — one move and one framing. "Slow push-in, medium shot." Written directions like "wide shot" or "dolly" are detected and applied.
4. **Light** — where it comes from and what it looks like. "Late afternoon sun through blinds, hard stripes on the wall."
5. **Mood** — two words at the end, never a paragraph. "Quiet, expectant."

Put together:

> A woman in a yellow raincoat crosses an empty parking lot in steady rain, puddles rippling around her boots, slow push-in from a low angle, sodium streetlights turning the wet asphalt orange against a blue-gray sky, lonely and calm.

Thirty-nine words. One subject, one action, one camera move, one light source, one mood. That is the whole trick; everything below is how people manage to break it.

## What breaks engines

The same failures show up on every engine in the roster, from Motion One to the premium ones:

- **Two subjects doing two things.** "A dog chases a ball while a man reads a newspaper" gives you a dog-shaped man reading a ball. Pick one subject; the other becomes background.
- **Sequences.** "She opens the door, then sits, then pours tea" is three shots. A clip is one. If you want three, use Film mode and write three scenes.
- **Abstract adjectives with no visual.** "Epic", "stunning", "cinematic" and "high quality" describe your hopes, not the frame. Replace each with a light or a lens.
- **Camera salad.** "Drone shot, close-up, tracking, orbit" makes the engine average four moves into a wobble. One move per clip.
- **Contradictory light.** "Golden hour with neon signs under a noon sun" is three lighting setups. Choose one.
- **Text and lettering.** Signs, subtitles and logos come out as alphabet soup. Describe the sign's shape and color; leave the words out.
- **Negatives inside the prompt.** "No rain, no crowd" tends to add rain and a crowd. Use the negative prompt field on engines that support it, and describe what *is* there instead.
- **Tag lists.** "masterpiece, 8k, ultra detailed, dramatic" is image-model dialect. Video engines want sentences.

## Before and after

Three prompts pulled from real drafts, and the rewrite that rendered.

**Before:** "Cyberpunk city, rain, neon, epic drone shot, very detailed, 4k."

**After:**

> A lone courier on a black motorcycle threads through slow night traffic on a rain-slick avenue, neon storefront signs smearing pink and cyan across the wet road, low tracking shot alongside the bike, headlights flaring in the mist, restless and electric.

The rewrite has a subject, a single action and a single move. "Epic" became a tracking shot; "neon" became specific colors on a specific surface.

**Before:** "A beautiful girl smiles at the camera, then walks away into a sunset, then looks back, cinematic."

**After:**

> A young woman with windblown auburn hair walks slowly away from the camera along a beach at sunset, glancing back once over her shoulder, static wide shot, contre-jour light rimming her silhouette in gold, wistful and warm.

One continuous action with a single beat inside it, one framing, one light direction. The "then, then, then" is gone.

**Before:** "Product shot of a perfume bottle, luxury, elegant, spinning, studio."

**After:**

> A faceted glass perfume bottle rotates slowly on a black mirrored surface, a single soft cross light catching each facet in turn, macro lens, shallow focus with the label softening in and out, tiny dust motes drifting in the beam, hushed and expensive.

"Luxury" became a mirrored surface and soft cross light; "spinning" became a rotation speed the engine can hold.

## One move, chosen on purpose

The studio exposes 17 camera moves as chips: Static, Slow pan, Push-in, Tracking, Orbit, Handheld, Close-up, Wide, Pull-back, Dolly zoom, Crane up, Crane down, Whip pan, FPV drone, Aerial pull-back, POV and Snorricam. Pick one per clip. If you prefer writing, the same words in the prompt are detected, so "orbit around the statue" does what the chip does.

A rough matching that has served us well:

- Stillness and portraiture: **Static** or **Slow pan**.
- Reveal or approach: **Push-in**, **Crane up**, **Pull-back**.
- Following a subject: **Tracking**, **Handheld** for nerves, **POV** for immersion.
- Scale: **Wide**, **FPV drone**, **Aerial pull-back**.
- Unease: **Dolly zoom**, **Snorricam**, **Whip pan** (sparingly; it eats half the clip).

More on each in [camera controls](/tools/camera-controls) and the [camera move](/glossary/camera-move) glossary entry.

## Length, format and engine choice

Clips run 5 or 10 seconds on [Motion One](/models/eroq-motion-one), the engine every account gets, and from 3 to 30 seconds across the premium roster depending on the engine and your plan. Free accounts cap at 10 seconds, Hobby at 15, Creator at 20, Studio and up at 30; the details are on [pricing](/pricing). Write for the length: a 5-second clip is one gesture, a 10-second clip is one gesture with a beginning and an end.

Formats are 16:9 wide, 9:16 vertical and 1:1 square. Vertical clips want a subject that stands; wide clips want a subject that moves sideways. Say which in the prompt ("she walks left to right") so the framing has something to follow.

If you already have the frame, skip half the prompt: drop in a reference photo or a Character and the engine animates it ([image-to-video](/tools/image-to-video)). The prompt then only needs motion, camera, light and mood; the subject is the picture.

## The free Enhance button

Next to the prompt box sits **Enhance / Surprise me**. It is a free prompt rewrite: paste a rough draft and it returns a director-grade paragraph in the structure above; leave the box empty and Surprise me invents one. It costs nothing, so the useful habit is to run it, read what it added, then edit. If it invented a camera move you did not want, swap the chip. If it made the light too specific, loosen it. The rewrite teaches the structure faster than any guide, including this one. Details on the [prompt enhancer](/tools/prompt-enhancer) page.

Every render lands in your library with its full recipe, so a prompt that worked is one Remix away from a variation. Failed renders refund themselves, so a prompt that did not work costs you a minute, not credits.

## FAQ

### How long should an AI video prompt be?

Between 30 and 80 words, as one paragraph. Shorter and the engine fills the gaps with defaults; longer and it starts dropping details, usually the ones you cared about most. Length is not quality. Specificity is.

### Should I describe the camera in words or use the chips?

Either. The chips are explicit and survive edits; written directions like "wide shot" or "dolly" are detected in the prompt and applied. Do not do both with different moves, or the engine gets two instructions.

### Does Enhance cost credits?

No. Enhance and Surprise me are free rewrites, and you can run them as many times as you like before you spend anything on the actual render.

Start with a working prompt and change one part at a time: [open the raincoat shot in the Video studio](/studio/video?prompt=A%20woman%20in%20a%20yellow%20raincoat%20crosses%20an%20empty%20parking%20lot%20in%20steady%20rain%2C%20puddles%20rippling%20around%20her%20boots%2C%20slow%20push-in%20from%20a%20low%20angle%2C%20sodium%20streetlights%20turning%20the%20wet%20asphalt%20orange%20against%20a%20blue-gray%20sky%2C%20lonely%20and%20calm.).
