The Unified Cinematic Video Prompt: One Block That Holds Every Shot Together
You generate six clips for a launch film. Individually, three of them are beautiful. Cut together, they look like six different brands shot on six different days by six different crews. The light flips sides. The grade warms then cools. The bottle is glossy in shot two and matte in shot five. Nothing is technically broken — and yet the sequence has no author.
That is shot-to-shot drift, and it is the single most expensive problem in AI motion work, because it is invisible until the edit. The fix is not a better model or a longer prompt. It is a unified cinematic video prompt: one fixed block of direction you reuse word-for-word across every clip, so the only thing that changes between shots is what the subject does.
This guide gives you the structure, a copy-ready template, a worked six-shot example, and the small habits that keep a sequence coherent from first frame to final cut.
Why your shots drift apart
A video model does not remember your last generation. Every run starts from nothing and fills in everything you did not state. Write “perfume bottle on marble, cinematic” six times and you will get six different interpretations of “cinematic” — because that word carries no fixed instruction. It is a mood, not a specification.
Drift therefore has a precise cause: under-specification. Each unstated variable is a coin the model flips on your behalf. Light direction, lens length, camera height, colour temperature, grain level, depth of field, motion speed — leave any of them open and it will change between takes, usually in ways your eye notices before your brain names them.
The counter-intuitive part is that more creativity in the prompt makes this worse. Rewriting each shot with fresh, vivid language feels like good craft, but every rewrite reshuffles the variables. Consistency comes from repetition, not invention. In a real production, the DP does not relight the set between shots. Your prompt block is that set.
The anatomy of a unified prompt
Split your prompt into two halves with completely different jobs. The variable half is one sentence and changes every shot. The fixed half is six slots and never changes inside a sequence — not a synonym, not a reordering, not a comma.
The variable line: subject and action
One sentence: who or what, doing one specific thing, with one emotional register. “The bottle rotates a quarter turn as a single drop traces the shoulder of the glass.” One action per shot. Two actions in one prompt is how you get mush — the model averages them instead of performing them.
Slot 1 — Camera and lens
State the shot size, lens length and camera height together: “medium close-up, 85mm, camera at subject height.” Lens length is the slot most people skip and it does more work than any adjective — it fixes compression, falloff and how the background sits behind the subject.
Slot 2 — Camera move
Name the move and its speed: “slow push-in, barely perceptible” or “locked-off, static frame.” If you leave this empty, most engines default to a drifting handheld wander that instantly reads as generated. In premium work, restraint is the tell of authorship.
Slot 3 — Light
Direction, quality, and a named source: “single soft key from camera-left through a diffusion scrim, deep unfilled shadow camera-right.” Direction is non-negotiable — a key that jumps sides between shots is the most jarring form of drift, and the one audiences feel even when they cannot articulate it.
Slot 4 — Palette and grade
Three colours maximum, plus the grade: “bone, warm amber, deep umber; desaturated editorial grade, lifted blacks.” Naming the blacks matters more than naming the highlights, because crushed versus lifted shadows is what separates a film look from a phone look.
Slot 5 — Texture and medium
“35mm film grain, subtle halation on speculars, shallow depth of field.” This slot is your fingerprint. It is also the cheapest way to unify footage generated by different engines, since grain and halation read as a shared capture medium even when the underlying renders differ.
Slot 6 — Negative space and mood
Close with the composition rule and the register: “generous negative space upper third, quiet and unhurried.” Composition instructions at the end tend to survive better than the same words buried mid-prompt.
The copy-ready template
Paste this into your notes, fill the fixed half once per project, then change only the first line per shot.
[ACTION] — one subject, one action, one emotional register. Camera: [shot size], [lens]mm, camera at [height]. Move: [move], [speed]. Light: [direction] [quality] key from [source], [shadow behaviour]. Palette: [colour one], [colour two], [colour three]; [grade], [black level]. Texture: [film stock] grain, [halation/bloom], [depth of field]. Frame: [negative space rule], [mood in two words].
Everything below the first line is your look lock. Write it once, save it, and treat editing it mid-sequence as a decision that costs you every clip you have already made — because it does.
A worked example: a six-shot fragrance film
Here is the same look lock carrying six different actions. Read the variable lines top to bottom and you have a shot list; read the fixed block once and you have the film’s visual grammar.
FIXED LOOK LOCK Camera: medium close-up, 85mm, camera at subject height. Move: slow push-in, barely perceptible. Light: single soft key from camera-left through a scrim, deep unfilled shadow camera-right. Palette: bone, warm amber, deep umber; desaturated editorial grade, lifted blacks. Texture: 35mm grain, subtle halation on speculars, shallow depth of field. Frame: generous negative space upper third, quiet and unhurried. SHOT 1 The bottle stands untouched as dust drifts through the key light. SHOT 2 A hand enters frame-right and rests fingertips against the glass. SHOT 3 The bottle rotates a quarter turn, a highlight travelling its shoulder. SHOT 4 The stopper lifts away, a thread of vapour following it upward. SHOT 5 A single drop falls and spreads across warm skin. SHOT 6 The bottle settles back, the room returning to stillness.
Notice what the action lines have in common: each is one movement, each is under fifteen words, and none of them mentions light, colour, or lens. That discipline is the whole method. The moment a shot line starts describing the grade, you have two competing specifications and the model will average them.
Notice too that shots 1 and 6 are near-mirrors. A unified prompt makes bookending effortless — the same lock plus a reversed action gives you an ending that rhymes with the opening without any manual matching in the edit.
Anchoring the lock to a still
A prompt fixes intent; a reference still fixes fact. The strongest workflow uses both: generate one hero frame you genuinely love, then feed that frame into every clip as the starting image while the unified prompt drives the motion. The still holds product geometry, typography and material behaviour that language alone cannot guarantee.
This is where image-to-video stops being a convenience and becomes a consistency tool. Practically: build the hero still in the Image Studio, save the look lock as a reusable visual signature, then move the frame into the Video Studio and run all six actions against it.
If you need a different angle for a given shot, do not re-prompt from scratch — regenerate the still from the same signature first, then animate. Change one variable at a time and drift stays traceable instead of mysterious.
Adapting the lock across engines
Different engines respond to the same words with different enthusiasm. The six slots stay; the emphasis moves.
Motion-forward engines need an explicit move verb or they invent one, so keep Slot 2 assertive and short. Reference-led engines weight the input frame heavily, so you can soften Slots 4 and 5 and let the still carry the grade. Photoreal engines respond well to physical-camera language — focal length, scrim, stop — while stylised engines respond better to medium language, like “hand-painted cel, held on twos.”
What never changes is the ordering. Keep the slots in the same sequence for every engine so that when a shot misses, you can see at a glance which slot got ignored and rewrite exactly that line. A consistent prompt shape is also a debugging tool.
Five habits that keep a sequence coherent
Lock before you scale. Generate two test clips with the full block before committing credits to the remaining four. If those two intercut cleanly, the lock is sound.
Never edit the lock mid-sequence. If it must change, re-run every shot. Half-updated sequences look worse than consistently imperfect ones.
Keep action lines under fifteen words. Long action lines steal weight from the fixed block and drift creeps back in.
Save the lock, not just the output. A look lock is a brand asset. Six months later it is what lets a new campaign match the old one without archaeology — store it alongside your prompt vault entries.
Judge in the timeline, not the preview. Clips look consistent alone and inconsistent in sequence. Always review cut together, at full size, in order.
When a unified prompt is the wrong tool
A single lock suits films with one visual world — product launches, mood pieces, editorial loops, atmosphere reels. It fights you when the brief genuinely needs contrast: a before-and-after, a day-to-night arc, a documentary intercut with archive texture.
In those cases, do not abandon the method — run two locks and switch deliberately at the story beat. Two intentional looks read as direction. Six accidental ones read as a mistake. That distinction, rather than any single model choice, is what makes AI motion work look commissioned instead of generated.
Frequently asked
- What is a unified cinematic video prompt?
- It is a single fixed prompt block — camera, lens, lighting, palette, grain and mood — that you reuse verbatim across every clip in a sequence, changing only the subject action line. The constant half is what keeps six separate generations feeling like one film.
- Why do my AI video shots look different from each other?
- Because each generation re-invents everything you left unspecified. If your prompt does not fix the light direction, lens length, colour grade and grain, the model picks new ones every run. Drift is not a model flaw; it is an under-specified prompt.
- How long should a cinematic video prompt be?
- Roughly 60 to 90 words: one action line, then a fixed block of camera, lighting, palette and texture. Longer prompts dilute the weighting on the details that actually carry the look, and short prompts leave too much for the model to guess.
- Does a unified prompt work across different video models?
- The structure travels, the wording does not. Keep the same six slots in the same order for every engine, then adjust phrasing — motion-heavy engines need an explicit camera-move verb, reference-led engines lean harder on the still you feed them.
- Should I write shot direction or emotion into the prompt?
- Both, but in separate slots. Emotion belongs in the action line where it changes per shot; direction belongs in the fixed block where it must never change. Mixing them is the most common reason a sequence loses its through-line halfway.
Build your hero still, save the look lock as a visual signature, and run every shot of the sequence against it without leaving the studio.