Seedance AI Video Prompting: The Director's Playbook
The real working manual behind professional Seedance prompts: motion hierarchy, camera language and continuity rules that turn AI video mush into cinema.
By Ezekiel 'Kiel' Orji — Co-founder, PxLabs · 2026-07-29 · 12 min read
Seedance AI Video Prompting: The Director's Playbook
This is not a tutorial — it is a working director's manual, paid for in failed generations, content-filter rejections and character-cap trims at three in the morning. By the end of it you will know why most AI video prompts render as mush, the exact hierarchy of what to write first in a Seedance prompt, and the camera and continuity rules that make a generation read as cinema instead of a hallucination. Whether you are learning how to make AI video ads, building an AI content workflow for your brand, or leading a marketing team into AI video, this is the foundation everything else sits on.
Key takeaways
- AI video is not a generation problem. It is a directing problem. The model is a department. You are the director. The prompt is your shot list, blocking notes, DP conversation and edit script in one block of text.
- The camera is the altar. If you do not direct the camera, you have not directed the film.
- Every shot must have one dominant motion and at most two supporting motions. More than that and the model averages — which means it renders mush.
- Models cannot render emotions, only the physical evidence of emotions. Translate every feeling into movement before you write it.
- The model defaults to perfection, and perfection is why AI video reads as AI. You have to put the micro-imperfections in the prompt.
- Beats do not exist in isolation. Chain them — dust, wind, light and momentum must carry from one beat to the next.
Why your AI video prompts produce mush
The shift from prompter to director is the shift that separates working AI video from amateur AI video. In the words of Ezekiel "Kiel" Orji, co-founder of PxLabs:
A prompter writes wishes. A director writes instructions. A prompter says "make it cinematic." A director says "35mm anamorphic, soft rim light from accretion disks at frame-right, camera locked, slight handheld breathing, deep blacks lifted slightly, fine 35mm film grain."
The prompter believes the model is creative. The director knows the model is a department that needs to be told what to do. The posture to adopt: when you sit down to write a prompt, you are sitting down with a competent crew who has never read your script and will execute exactly what you specify and nothing more. Every choice you do not make, they will make for you, badly. Every choice you do make, they will respect.
This is The Director, Not the Prompter — the first principle, and the one everything in this AI video prompting guide flows from.
Two companion principles complete the philosophy. The camera is the altar: the camera is the only thing the audience ever sees, so if you do not direct the camera, you have prompted a hallucination, not directed a film. And the world is the star: AI models render foreground subjects confidently but struggle with the inhabited specificity of a world — so spend your prompt budget on the world: the materials, the light, the dust in the air, the small inhabitants at frame edges, the wear on the surfaces. That is what AI is worst at and what cinema needs most.
The motion hierarchy: what to write first in a video prompt
This is the single most important technical principle in the whole practice. Every shot must have one dominant motion and at most two supporting motions. More than that and the model averages, which means it renders mush.
The hierarchy, in priority order:
- Camera motion — the most important. The camera's movement IS the audience's perceptual experience. A locked tripod is a choice. A slow forward push is a choice. Choose one.
- Character gesture — what the protagonist does. A turn of the head. A blink. A hand rising. Pick the single gesture that earns the beat.
- Jewelry, fabric, or accessory secondary motion — the small life signals. Earrings swinging. A robe drifting. A scarf catching wind. These prove the world is real.
- Environmental movement — dust, wind, water, smoke. Atmospheric breath that never competes with the protagonist.
- VFX accents — flares, particles, light shifts. The lowest priority. If VFX dominates, the shot is failing.
Write your prompt in this order. Camera first. Always. If you cannot name what the camera is doing in one sentence, you have not finished thinking yet.
The One Dominant Action Rule
Each beat shows ONE thing happening. Not three. Not five. One.
If your beat description contains the word "and" connecting actions, you have probably overloaded the beat. Compare:
Overloaded: "She turns toward the lens AND her earrings swing AND the dust drifts past AND a figure passes in the background" — a beat that will render as four blurred half-actions.
Written correctly: "She turns slowly toward the lens. (Secondary: earrings swing once. Background: dust drifting in light shafts. Frame-left: a figure passes.)"
Same content, but now the model knows the hierarchy. The principle: one ACTION per beat, surrounded by smaller LIFE SIGNALS. Action is what we are looking at. Life signals are what makes us believe the world is alive while we look at the action.
Physical continuity over abstract emotion
AI video models cannot render emotions. They can only render the PHYSICAL EVIDENCE of emotions.
Bad prompt: "She feels grief." The model has no idea what to put on screen.
Good prompt: "Her eyes lower. Her shoulders drop one centimeter. A single tear catches the light at the corner of her eye but does not fall. Her hand rises halfway to her face and stops."
Now the model has physical evidence of grief, and the audience reads grief from the evidence — exactly the way emotion works on real film sets, where actors and cinematographers build feeling from physical detail first. Translate every emotional intention into physical movement before you write it into a prompt. What does not work: "she looks sad," "emotional," "melancholic," "cinematic and moving." If you cannot name the physical evidence, the model cannot render it.
Micro-imperfection theory: why AI video looks like AI
The model defaults to perfection. Perfect lighting. Perfect motion. Perfect symmetry. Perfect surfaces. This is why AI video reads as AI. Real cinema is full of micro-imperfections — the camera shakes slightly, the actor's hair is not quite in place, there is dust on the lens, a piece of furniture is rotated three degrees off square.
You have to put the imperfections IN the prompt. They will not appear by accident. Phrases that work:
- "slight handheld breathing"
- "anamorphic lens breathing"
- "rain-flecks on the lens glass"
- "missed focus, briefly catching the railing edge"
- "slightly rotated chairs, compressed cushions"
- "uneven object placement"
- "the camera makes mistakes"
The single most powerful instruction in the manual: "camera makes mistakes — missed focus, wrong framing, briefly catching the railing edge. THIS person's footage, recovered." That one line turns AI generation into found-footage cinema.
How Seedance actually reads your prompt
This is the engineering knowledge that most Seedance prompts ignore.
Seedance is verb-driven. The first action verb in a beat is the one it renders hardest; subsequent verbs are weighted lower. "The camera arcs slowly around her" — Seedance commits to the arc. Add "as she turns her head while her earrings swing" and the model executes the arc, weakens the turn, and probably misses the swing. The fix: put the most important verb first, and give supporting verbs their own clauses with their own subject — "The camera arcs slowly around her. (Secondary: she turns her head. Tertiary: earrings swing once.)"
Seedance rewards technical camera vocabulary. It understands specific lens lengths ("35mm anamorphic", "100mm macro", "14mm ultra-wide"), specific moves ("slow forward push", "dolly", "orbit", "crash zoom"), and cinematographer references that load whole lighting styles. Do not write "the camera moves." Write "27mm cathedral wide, deep depth of field, slight perspective compression at distance." In Kiel's own account, the single biggest jump in quality came the day he started naming lens focal lengths in every beat.
Seedance hates ambiguity. It resolves every unspecified degree of freedom by defaulting to the most generic reading in its training data. Always specify camera position, light source, character orientation and time of day. Every degree of freedom you leave the model is a degree of freedom you have lost.
Density has a breakpoint. Roughly one fully-described visual element per second of generation — about 10 named elements max in a 10-second prompt. Past that you get the classic over-density failures: motion smearing, identity collapse, colors averaging toward grey, light flattening, textures going generic. The cure is never to add more instruction. The cure is to cut.
Spatial layering works — with spatial keys. "Foreground: drifting sand across the obsidian floor. Midground: her seated form, head turned camera-left. Background: red desert mesas under teal sky through broken openings." Seedance reads foreground/midground/background as spatial keys and renders each layer separately. Write "sand drifts and she sits and the desert is visible" and all three elements compete for the same plane.
Know what breaks it. More than 3 named characters in one beat (identity collapses across them). Conflicting motion instructions. Hyper-specific in-frame text. Time-of-day reversals mid-prompt. Hex codes for color — use color names. More than one cinematographer reference per prompt.
A prompt structure that works: the Cinematic Formula
The workhorse structure behind 90% of the professional prompts in the manual:
SCENE → SUBJECT → MOTION → CAMERA → LIGHTING → ATMOSPHERE → CONSTRAINTS
Where are we, who is in frame (locked identity description), what happens (in hierarchy: camera, character, environment), specific lens and movement, then lighting specified as source + color temperature + quality + behavior — "warm lighting" is a tragic underuse of the model; "single key from upper-left, tungsten 3200K, soft volumetric haze, reflected light shifts subtly across gold leaf" is directing. Atmosphere carries the mood register, and constraints — the negative prompts and identity locks — go last, because the last thing the model reads is the most recently considered. This formula scales from 1,000-character prompts to 4,000-character prompts.
For quick tests, the Minimal Formula gets you 80% of the way in four clauses — SUBJECT → ACTION → CAMERA → ENVIRONMENT:
A Yoruba woman in a black gele turns slowly toward the lens. Locked tripod, 35mm anamorphic, eye-level. Brown void background, soft directional key light from upper-left.
For multi-beat sequences, timestamp every beat ([0:00-0:03]) — Seedance treats timestamps as hard structural constraints and uses them to allocate motion budget. Avoid vague temporal language like "first" and "then"; the model treats those as suggestions.
Camera language: the moves and when to use them
The camera is not a recording device. It is a character, and its movement is emotional storytelling:
- Locked tripod — this is bigger than the audience can intervene in. Witness it. Don't use when you need intimacy.
- Slow forward push — come closer. There is something worth seeing. The most reverent move in fashion-film vocabulary.
- Slow pullback — step back. The full scale is more than you knew. The trailer's payoff move; it earns the title card.
- Handheld — this is real. This is a body holding a camera. Never on monumental subjects; it undercuts gravity.
- Dutch tilt / orbital — something is wrong, or this person is a monument worth circling. Never orbit a subject who is also moving — two simultaneous motions confuse the model.
- Crash zoom — shock, discovery. Once per piece maximum. Multiple crash zooms in 15 seconds read as panic.
- Macro insert — a single detail seen with reverence, sandwiched between wides: wide, macro, wide.
And choose focal length for the emotional effect, not just the framing: 14mm ultra-wide for awe and cosmic scale, 35mm anamorphic as the cinematic default, 85mm for emotional close-ups, 100mm macro for texture beats. A practical bonus for anyone building an AI content workflow: long lenses (85mm+) tend to produce more flattering character renders, because the model has more training data for portrait-lens portraits than wide-angle face shots.
Environmental chaining: how professionals keep AI video consistent
This is the section that separates amateur AI prompting from professional AI prompting. Most prompts write each beat as a standalone shot. Pro prompts chain them: the dust that was drifting in shot one should still be drifting in shot two; the light that was warm in beat three should still be warm in beat four unless something motivated the change.
The chaining checklist for every beat transition: Where did the dust go? Where did the wind go? What is the light doing now compared to a moment ago? What momentum is carrying over?
Dust is the single most useful continuity tool in AI video — models render it beautifully and it signals atmospheric reality without competing for attention:
[0:00-0:03] Fine sand drifts across the floor in the breeze. [0:03-0:06] Macro on the glyph — dust still drifting across the polished surface. [0:06-0:09] Wide shot — same dust suspended in the volumetric light shafts.
Same dust, different beats — the audience reads continuity without consciously noticing it. Wind gets specified once, globally: "GLOBAL WIND: gentle drift from frame-right to frame-left throughout. All fabric, dust, and particle motion follows this direction." Fabric carries continuity through decay states — drifting, still settling, at rest but microvibrating.
The most actionable trick of all is the carryover sentence: at the end of every beat description, write one sentence naming what carries into the next beat — "Carries forward: amber glyph light continues, robotic hand still extended." You can strip these before submission to save characters, but writing them forces you to think in continuity rather than snapshots.
Learn this live
Reading a director's manual is one thing; generating alongside a director is another. This material is taught hands-on in our Generative Cinematography course at /classes/generative-cinematography — take Synthography 101 first if you are new to generative imagery. It is also the kind of thing pulled apart daily in the PxLabs community at /community, the home base for AI training in Lagos and one of the most active AI communities in Nigeria for people learning AI video production.
And if you are a brand that wants this made for you rather than by you — AI video ads, product films, campaign work — that is what our studio does. Start at /contact.
Next in the series
Part 2 of Director's Prompt: The AI Secret Sauce covers the question every team asks next: Seedance vs Runway vs Dreamina — which model to choose for which job, where each one wins, and how the same director's brief translates across all three.