A 30-second continuous take is a directing problem before it’s a model problem. We covered that side when Wan 3.0 landed: what the model ships with, and the pre-production that makes a long take worth generating.
What’s less settled is the part in between: the actual words you type. Wan 3.0 reads a prompt in a specific way, with formulas for different jobs, a syntax for naming references, and controls that only exist as literal phrases you have to know to write. Miss them and the model fills the gaps for you, usually with cuts you didn’t ask for, dialogue you didn’t write, and music you didn’t want.
This is the language side of it. A Wan 3.0 prompt guide: 7 formulas, the reference syntax, how to direct sound, how to edit what you already made, and 5 prompts you can run as they are.
How Wan 3.0 reads a prompt
A single sentence will get you a complete video. Precision is what gets you the video you actually pictured.
The model treats a prompt as a set of decisions rather than a description, and anything you leave undecided it decides for you. That matters more here than on a short-clip model, because 30 seconds contains more decisions: how many shots, where the cuts fall, who speaks, whether there’s music at all.
Three controls that only work if you write them
Wan 3.0 dropped the negative prompt field that 2.7 carried. There’s no separate box for exclusions, so anything you want kept out goes in the prompt body as a plain instruction.
- Shot count. Write Shot 1 [0-3s]: … Shot 2 [3-6s]: … to structure a sequence. For an unbroken shot, write One continuous take. Say nothing and the model cuts wherever it likes.
- Dialogue. Leave it out and characters will start talking. For silence, write No dialogue throughout the entire video.
- Music. Leave it out and the piece gets scored. To stop that, write No background music.
Those three lines account for most of the distance between a generation that looks close and one you can use.
Prompt length is not the constraint
You have up to 20,000 characters. The prompts in this guide run long on purpose: every clause is a decision taken off the model’s hands. Length isn’t the goal, but running out of room won’t be your problem.
The 7 formulas
Each one covers a different job. Learn which is which and you stop writing the same prompt for every task.
| Formula | Structure | Use it for |
|---|---|---|
| Basic | Subject + scene + motion | Exploring an idea, first drafts |
| Advanced | Subject + setting + motion + aesthetic control + stylization + sound | Full control over a finished shot |
| From a start image | Motion description + camera movement | The frame already fixes subject and style |
| Sound | Voice + sound effects + background music | Anything where audio carries weight |
| Reference | @Reference subject + action + dialogue | Characters, products, or locations that must hold |
| Multi-shot | Overall description + shot number + timestamp + shot content | Sequences and narrative |
| Editing | Editing target + editing action | Changing a video you already generated |
Aesthetic control is its own vocabulary: light source, lighting environment, shot size, camera angle, lens, and camera movement. Naming all 6 gets you further than any amount of adjectives.
Stylization is the look, stated plainly. Cyberpunk, line-drawing illustration, post-apocalyptic, 16mm documentary, claymation. Style holds without bleeding between looks, so a multi-style project stays separated.
The formulas stack. A multi-shot narrative with two referenced characters and directed audio uses three of them at once, which is what most of the prompts further down are doing.
Reference syntax
References are cited by name and by order: @img1, @vid2, @audio1. The numbering follows the order you loaded them.
Bind, then bound
The most common failure is uploading references and assuming the model will work out what each one is for. It won’t. A character sheet arrives carrying a face, a build, wardrobe, and usually a pose and a background you have no use for. State what each reference defines, and state what it should ignore.

@img1 — character reference

@img2 — environment reference
Prompt
@img1 defines the skater’s face, build, and wardrobe. Do not use the pose or the background from the reference sheet. @img2 defines the environment: the stairs, the metal rail, and the graffiti wall.
The exclusions are doing as much work as the definitions. Since there’s no negative prompt field, that’s where they live.
Then modify freely on top. Wardrobe changes, damage, weather, and age all layer over a scoped reference without breaking the identity underneath.
Cite the same asset more than once
You can mention a reference at several points in a prompt to control where it applies, shot by shot. This is how you keep a character consistent through a sequence without restating their whole description each time.
What each media type contributes
- Images anchor faces, products, and locations
- Video contributes timing and camera grammar, not content
- Audio carries voice timbre, so a character speaks in a specific voice rather than a generic one
Here’s what came back.
Two references, one prompt. The character sheet scoped to face, build, and wardrobe; the location plate carrying the plaza. The skater holds through 7 shots and as many camera moves, from a ground-level dolly to a 180-degree orbit to a closing crane, without becoming someone else. Consistency across cuts is the harder version of the problem, and it’s the one that matters on a sequence.
One structural rule
Start and end images can’t be combined with references. Load one and the other switches off. Magnific tells you which is active the moment you upload, but it’s worth knowing before you plan a shot around both.
Directing sound
Audio generates in the same pass as the picture, which means it’s directed in the same prompt. Most people write the picture and let the sound happen. It’s the cheapest quality you’re leaving behind.
A butcher shop in Bogotá, late at night. 30 seconds, one continuous take, sound generated with the picture. What follows are the audio instructions from that prompt, lifted out of a much longer brief that also carries blocking, camera, references, and style. The sound is a fraction of the text, and it’s the fraction most people skip.
The voice recipe
The line + emotion + intonation + speech rate + timbre + accent
6 parameters. Writing only the line gets you a generic read. This is one 6-second stretch of the butcher shop prompt, with the voice direction sitting inside the action:
Prompt
11–17s: ELIÁN stops wiping. Camera settles into a loose two-shot, both men in frame, deep space between them. ELIÁN raises his head for the first time: {¿A avisarme qué?} MAURICIO holds the look, then breaks it: {Que dieron tu nombre.} ELIÁN sets the knife down on the steel. The sound is loud.
2 lines, 2 different instructions, and neither of them is a tone set once at the top of the prompt.
The accent is declared separately, once, and it governs the whole scene: Dialogue language: Spanish, Colombian Bogotá accent (rolo). Neutral and clipped, soft consonants, no sing-song intonation.
Accent is worth pushing on specifically. Name the language and the regional accent and the delivery follows, which is what makes one generation viable for a specific market instead of a neutral one. Describing what an accent isn’t helps too: a Bogotá accent lands better when you also say clipped and without sing-song intonation, because that rules out the Colombian accent most models default to.
The full scene runs 6 exchanges this way, each one directed where it happens. The last one, the line carrying the worst news, is marked no weight on it at all, and it only lands because it’s written that way. Set a single tone at the start and the model will escalate into the reveal instead of dropping it.
Sound effects
Written as a single instruction, the formula looks like this:
Source object + action + ambient sound
Prompt
A boning knife is set down on a stainless steel counter, one hard clack, in a quiet tiled room at night.
Inside a full scene, though, the three parts split up. The ambient half gets established once, in the location block, where it doubles as visual description:
Prompt
Location: interior of a small butcher shop at night. Hard overhead fluorescent, one tube flickering faintly. Cold-room fog drifting low across a wet tile floor. A plastic strip curtain hangs in the mid-ground. Stainless steel counter, hanging steel hooks, rain outside the front window. Practical light only.
Not one word of that is written as audio direction, and 3 of those details carry the entire sound bed: the flickering tube, the wet tile, the rain. The effect itself then gets 4 words, dropped where it happens:
Prompt
ELIÁN sets the knife down on the steel. The sound is loud.
That lands because the quiet around it was built 20 lines earlier. The same knife reads as nothing in a busy room and as a full stop in a silent one, and it’s the location block that decides which.
Music
Background music + style
Or its absence, which is a decision worth making deliberately. The butcher shop scene ends with No music. Just diegetic sound. That’s 3 seconds of typing, and it’s what leaves room for the fluorescent hum and the knife to carry the tension instead of a score doing it for them.
Editing and extending
The capability with no real equivalent elsewhere, and the one almost nobody is using. On a video you’ve already generated, you name what changes and everything else holds.
The formula
Editing target + editing action
And the clause that makes it work: keep everything else the same. Leave it out and the model takes liberties in places you never mentioned.
Prompt
In @vid1 the car changes to 90s classic model and to red color, while its movement and the rest of the video remain unchanged.
What you can change
- Elements — objects in frame, added, replaced, or removed. Name the object and the action, nothing more: add a boulder on the beach, replace the cat with a dog, remove the woman’s sunglasses.
- Lighting — relight the scene or kill one specific shadow. Brighten the overall lighting, soft three-point lighting, remove the hard shadow on the right side of his face.
- Style, color, and weather — a property of the whole piece rather than a thing inside it. Change the visual style to claymation, change to a warm color palette, change the video to an overcast day.
- Motion and position — where subjects sit in frame and how they move through it. Swap the positions, move the ostrich to the front-left of the frame, keep consistent walking movements and pace.
- Dialogue — a new line in the same voice. Always add maintain the character’s voice timbre and tone, or you may get a different voice back.
- Plot — the most aggressive version. You’re changing what happens, not how it looks. Reshape the storyline of @vid1: have the man pick up the guitar and start playing.
- Wardrobe and props from a reference — the video supplies the motion, an image supplies the what. Have the man wear the hat from @img1 and the shirt from @img2.
That clause belongs on every one of them.
Extending
Extension is not editing. Content, style, and composition are already fixed by the source video, so the prompt describes only what happens next.
- Forward: Extend @vid1 by 15 seconds: …
- Backward: Extend @vid1 backward by 15 seconds: …
- Both, as an inserted middle segment: Extend 2 seconds backward and 3 seconds forward as the middle segment: …
Extension prompts run long on action and dialogue and short on description, because the look is inherited.
Prompt
Extend @vid1 by 10 seconds, where the car moves forward and the camera follows it from behind
5 prompts to run
Copy them, run them, break them, rewrite them.
1. A revision on delivered work
Prompt
Edit @vid1: change the model’s jacket to the cream linen blazer from @img1 and warm the overall grade half a stop. Her movements, the camera move, the background, and everything else in the frame stay exactly the same.
2. One product, three setups
Prompt
@img1 defines the bottle’s structure, label artwork, and colors. Do not use the image background. Generate a 15-second product film, 9:16. Shot 1 [0-5s]: the bottle on polished dark stone, a single hard key from upper left, the camera orbiting 90 degrees at object height. Shot 2 [5-10s]: macro across the label and the cap seam, light sweeping slowly. Shot 3 [10-15s]: pull back to a wide, the set lighting up fully, the bottle centered and still. Maintain consistency: structure, label artwork, materials, and color identical to @img1 in every shot. Audio: a low sustained tone, one clean glass click as the orbit completes. No dialogue throughout the entire video. No background music.
3. The same scene, cast for one market
Prompt
Define the woman in @img1 as Subject 1. Her dialogue voice references @audio1. A bright open-plan office, late afternoon, floor-to-ceiling windows behind her. Medium shot, handheld with minimal drift. Subject 1 looks up from her screen, half-smiles, and speaks straight to camera. Dialogue language: Spanish, Colombian Bogotá accent, warm and unhurried: {Llevo tres semanas con esto y ya no vuelvo atrás.} She shrugs and lands the close, lower, almost confidential: {Ya está. Ese es el truco.} Maintain consistency: same face, same hair, same wardrobe, same room throughout. Style: natural window light, no aggressive grade, slight exposure imperfection. Audio: room tone, distant traffic. No background music.
4. A brand film in one continuous take
Prompt
A 30-second single continuous take, no cuts, through a working ceramics studio at first light. Warm, tactile, unhurried. The camera enters low at the doorway and moves forward through the room in one unbroken push. 0-8s: past drying racks stacked with unglazed bowls, dust suspended in the window light. 8-16s: the camera rises to counter height and arcs around a potter’s hands centering clay on the wheel, water running over the rim. 16-24s: it continues past the wheel to a kiln door swinging open, heat distorting the air above it. 24-30s: it pulls back to a wide of the whole studio as the room fills with morning light. Style: 35mm warmth, shallow depth of field, amber and deep brown palette, film grain. No on-screen text. Audio: the wheel motor, running water, the kiln door, a low acoustic guitar line underneath. No dialogue throughout the entire video.
5. A storyboard becomes a sequence
Prompt
Following the storyboard in @img1. Shot 1: wide shot of a small train station platform on a rainy night, a long-haired girl standing alone holding a blue umbrella, the station house behind her emitting warm light. Shot 2: medium shot, the girl and a short-haired boy with a schoolbag face to face, pouring rain, railway tracks extending between them. Shot 3: close-up of her face through the transparent umbrella with raindrops sliding across it, eyes slightly red, softly saying {Remember to send a message when you get there}. Shot 4: over-the-shoulder from behind the boy, a train with glowing headlights approaching in the distance, the halo diffusing through the rain. Shot 5: close-up of hands, a folded note passing into his hand, her fingertips trembling. Shot 6: the girl alone at the platform edge, the boy gone, the rain easing off.
Try these prompts with Wan 3.0
Prompting FAQ
Does Wan 3.0 have a negative prompt?
No. Version 3.0 removed the negative prompt field and prompt extension. Exclusions go inside the prompt body as plain instructions, the same way the shot, dialogue, and music controls do.
How do I stop it from cutting my shot into pieces?
Write One continuous take or Generate a single shot. Without that line, the model decides shot count on its own based on the story it reads in your prompt.
How do I stop it from inventing dialogue?
Write No dialogue throughout the entire video. If you do want dialogue, write the exact lines and the model will keep them.
How do I reference more than one character?
Cite them separately by name and order, then define each one: @img1 as Subject 1, @img2 as Subject 2. Mention each again at the points in the prompt where their behavior matters.
Can I change the voice without regenerating the video?
Yes, that’s a temporal edit. Ask for the new line and explicitly ask to preserve the original voice timbre and tone in the same instruction.
How long should a prompt be?
As long as the number of decisions you want to make. The ceiling is 20,000 characters, which in practice means length is never the limiting factor.
Do the formulas have to be used one at a time?
No. Most real prompts stack them. A referenced character delivering directed dialogue across a timestamped sequence is using three formulas at once.
Your turn
Take one of the 5, change a single variable, and watch what moves.