blog/guides·Sep 11, 2026·6 min·by the eroq team
Flicker and morphing in AI video — why faces drift mid-shot
Clothing that changes colour halfway, faces that slide, backgrounds that churn. What temporal instability is, what reduces it and what does nothing.
The clip plays and for two seconds it is perfect. Then the jacket goes from navy to black, the logo on the wall rearranges itself into a different logo, and somewhere around second four your character's face becomes a close relative of your character. Nothing failed. This is temporal instability, and it is the most common reason a technically successful render is still unusable.
The model is not watching its own footage
A video engine generates frames conditioned on what came before and on your prompt. What it does not do is hold a persistent record of the scene — there is no variable anywhere saying this jacket is navy or this is the same person as three seconds ago. Consistency across frames is an emergent property of the model being good at continuation, not a guarantee it is enforcing.
So drift is the default and coherence is the achievement. Everything that reduces flicker works by shrinking the window in which the model can wander, or by pinning down what it starts from.
Where it shows up first
Instability is not evenly distributed. It attacks the same things in the same order, which makes it predictable:
- Small faces. A face occupying a tenth of the frame gets a tenth of the attention and reconstructs differently every few frames.
- Hands and fingers, for the same reason they break in stills — plus motion.
- Printed text and logos. No current engine holds legible text across a shot. Letters rearrange. Treat any text in frame as guaranteed to churn.
- Patterned fabric. Stripes, checks and prints re-roll constantly. Plain clothing is dramatically more stable than a herringbone coat.
- Crowds and secondary characters. Anyone the model is not focused on will morph freely in the background.
- Churning backgrounds — foliage, rain, water, fire, smoke. These look like motion to the engine and it will happily reinvent them each frame.
If you are seeing morphing, check that list before you touch the prompt. Half of all flicker complaints are a patterned shirt or a sign on a wall.
What actually reduces it
Five changes, in descending order of how much they help.
Render shorter. Drift compounds, so a five-second clip has less room to wander than a ten. If a shot morphs at ten seconds, the same shot at five frequently holds. Two clean five-second clips cut together also read better than one wobbly ten.
Start from a frame. Image-to-video pins the first frame, so the model begins from a fixed, approved picture instead of inventing the subject and then failing to remember it. This is the single biggest lever, and the workflow is simple — render the still, approve it, hit Animate on it in your Library.
Simplify the background. A wall, a sky, a shallow depth of field. Anything you blur out is something that cannot churn, which is what "f/1.4" in a prompt is really buying you here.
One subject. Two people is more than twice as hard as one. The engine's attention splits and the one it is not looking at deforms.
One camera move, slowly. Fast moves reveal new geometry constantly, and new geometry is invented geometry. A slow push-in is the most stable thing you can shoot.
A prompt built on all five, ready to paste:
A man in a plain grey sweater sits at a bare kitchen table and looks up toward the window, holding still; slow push-in from a medium shot, soft window light from the left, plain wall behind him, shallow depth of field, a quiet and patient mood.
Plain fabric, one subject, one action, one slow move, nothing in the background to reinvent.
A seed makes your comparisons honest, not your clip stable
All three Seedance engines honor a seed — the same seed with the same prompt produces the same take. This is easy to misread as a stability feature. It is not.
A seed does not reduce morphing within a shot. It makes the shot repeatable, which is what lets you change one word and know that the difference you are looking at came from that word rather than from the dice. That is enormously useful when you are diagnosing flicker and useless as a cure for it. One practical consequence worth remembering: a batch of four takes with a fixed seed gives you four identical clips, so unset the seed when you actually want variety. Seeds and the remix workflow has the longer version.
For a shot that has to end somewhere specific — a loop, a reveal, a transition — the Seedance family also takes a first and last frame pair and interpolates between them. Both ends pinned is the most constrained a clip can be here, and constrained is stable.
What does not help
Adjectives about consistency. "Consistent face, no morphing, stable, coherent" are words about the output, not about the picture. The model has no mechanism for them and they dilute the words that do steer it.
A negative prompt, on most engines. This one is worth stating flatly because the field is accepted everywhere: the video negative prompt is native only on Kling 2.5 Turbo. On Motion One, the Seedance engines, Hailuo and Veo it is taken by the API and then quietly dropped before the engine sees it. If you want "no distorted faces" to do something in video, you have to be on Kling.
Higher resolution. 1080p renders the morph in more detail. That is all it does.
Holding a look across a whole sequence
Within-shot stability is one problem; keeping a character recognizable across twelve shots is another, and it has a different answer — a Character with reference photos, reused in every scene, so each clip starts from the same face rather than from the same sentence. Consistent characters in AI video covers that end of it properly.
The two stack well. Anchor the identity with a Character, anchor the frame with an image-to-video start, keep the clips short, and cut. Most footage that looks stable was assembled that way rather than rendered in one take.
Open the video studio and try your morphing shot at half the length with a start frame.
FAQ
Why does clothing change colour halfway through a clip?
Because nothing in the engine stores the colour — each frame is re-predicted, and a garment with a pattern or an ambiguous shade drifts fastest. Plain, high-contrast clothing survives a shot far better than stripes or prints. Starting from a reference frame also pins the colour at frame one.
Will a seed stop my video from morphing?
No. A seed makes a take reproducible so you can compare two versions fairly, but it does not make the frames inside that take more consistent with each other. It is a diagnostic tool, and it is only honored on the Seedance engines. Shorter clips and a start frame are what actually reduce drift.
Can I keep text or a logo readable in an AI clip?
Not reliably on any engine here. Letters are reconstructed per frame and will rearrange, so plan on adding titles, logos and lower thirds in your editor after the render rather than asking for them in the prompt. Leave the space for them in the shot instead.
Why does the negative prompt not do anything in my video?
Because it is only a native control on Kling 2.5 Turbo. The other video engines accept the field and ignore it, so what you typed never reaches the model. Either move the shot to Kling, or put the effort into a shorter clip, one subject and a simpler background.
Make this with the models behind the post — start with 50 free credits , or browse every engine and its price .