Build the core prompt
Seedance 2.5 accepts natural-language direction. Begin with the visible event, then add scene, look, camera, and sound only where those details help control the result.
- Subject and event
- Name who or what is on screen and what visibly happens.
- Scene and environment
- Set location, time, weather, spatial relationships, and background state.
- Visual treatment
- Describe lighting, palette, material, texture, or overall mood.
- Camera and cuts
- Specify shot size, angle, movement, focus, or how one shot becomes another.
- Audio
- Direct dialogue, voice quality, ambience, effects, or music.
<Subject> performs <main action or event> in <scene and environment>.
The image uses <visual style>.
Frame the event with <shot size, angle, camera movement, or cuts>.
Audio contains <dialogue, ambience, sound effects, or music>.A ceramic artist finishes a pale blue cup in a pottery studio at dawn, lifts it from the wheel, and sets it at the center of a wooden shelf.
Soft window light reveals a fine sheen on the wet clay; the workbench stays clean and orderly.
Begin on a medium view of the wheel, push slowly toward the cup's surface, then cut to a frontal shelf composition.
Keep the quiet wheel hum, clay friction, and subtle room ambience.Expand a short idea in four passes
A production prompt is easier to read when it moves from context to story, then closes with rules that apply to the whole clip.
- 01Assign the referencesIf files are attached, identify each one and state whether it controls a character, place, movement, voice, or another specific attribute.
- 02Write the creative briefSummarize the subject, location, central event, visual direction, and any distinctive camera behavior in one sentence.
- 03Describe the progressionUse a storyline or time ranges. For each beat, cover the visible action and add camera, dialogue, effects, or exclusions only when they matter.
- 04Finish with global rulesRestate continuity requirements that must hold throughout, plus unwanted elements such as extra subtitles, dialogue, or background music.
Leave out any block you do not need. Resolution, duration, aspect ratio, and other selectable generation settings belong in the UI or API rather than inside the prompt.
Choose reference materials deliberately
Seedance 2.5 can combine as many as 50 reference assets. The hard input ceilings and the stability-oriented recommendations are not the same thing: you can go beyond the recommended range, but the result may become harder to control.
| Material | Input limit | Recommended working range |
|---|---|---|
| Images | Submit a maximum of 30 images, with a 4K cap on each image | For a practical working set, use 1–8 unique subjects across the subject-reference images |
| Videos | Provide as many as 10 video clips, totaling 30 seconds or less | Aim for 1–5 unique subjects, with 5–10 seconds of footage for each subject |
| Audio | Include no more than 10 audio clips and keep their total runtime within 30 seconds | Use only the task-relevant dialogue, vocal qualities, ambient sound, or music |
| Video editing | A single source video can be supplied alongside reference images | The suggested setup is a source shorter than 20 seconds plus 1–5 reference images |
Larger sets can still work—for example, 9–12 subjects in image references, 6–10 subjects in audio/video references, or 6–8 images for an editing task—but stability can decline as the set expands. When more than five subjects need several views, provide separate view images instead of merging all angles into one collage.
Define exactly what each reference contributes
Do not ask the model to infer which person, prop, place, movement, or sound an upload represents. Bind each asset in the prompt and add exclusions whenever an unwanted background, person, or composition might leak into the result.
@Image 1 defines <subject>'s appearance, clothing, structure, or material>. Do not use <irrelevant content>.
@Video 1 defines <motion, camera path, or pacing>. Do not use <identity, wardrobe, or setting> from the video.
@Audio 1 defines <speaker or sound type>'s voice, dialogue, ambience, or music.
<Subject> completes <main action or event> in <scene>.
Use <visual style> and <camera treatment>.@Image 1 defines the ceramic artist's face, hairstyle, and dark green apron. Ignore its background.
@Image 2 defines the wooden workbench, window position, and dawn lighting of the pottery studio. Ignore every person in that image.
@Video 1 supplies only the rhythm of shaping clay, lifting the cup, and setting it down. Do not inherit the demonstrator's identity, clothing, or location.
The ceramic artist completes a pale blue cup, lifts it from the wheel, and places it in the middle of the shelf. Start with a medium shot and move slowly toward the clay texture. Retain wheel, clay, and room sounds.Multiple views of one object
When several images are angles of the same person or product, say so explicitly and lock the expected object count.
@Image 1 defines the front view of one folding desk lamp.
@Image 2 defines the left-side structure of that same lamp.
@Image 3 defines the right-side structure of that same lamp.
@Image 4 defines the rear structure of that same lamp.
All four images describe a single folding desk lamp. Keep exactly one lamp in the video.If a reference video already expresses the action, camera path, and timing well, describe only what should be inherited. Repeating the motion in different words can introduce conflict. A blockout clip usually supplies spatial structure and movement, so you still need to define the final characters, setting, action, and look.
Direct audio, dialogue, and visible text
Natural language is enough for most prompts. Use the syntax below when music, effects, spoken lines, and subtitles need to be separated clearly.
| Content | Syntax | Example |
|---|---|---|
| Music | ( ) | (Soft rhythmic piano plays underneath the scene) |
| Sound effects | < > | <A bicycle bell rings in the distance> |
| Dialogue | { } | {I knew you would come back.} |
| Subtitles | 【 】 | 【Chapter One: The Departure】 |
Reinforce the spoken language
Put the language before the line when it is not Chinese. If the system still chooses the wrong language—or a regional voice matters—state language, variety or accent, delivery, speaker, and dialogue in that order.
The girl says softly in Japanese: {もうすぐ着きます}
Dialogue language: American English. The woman speaks in relaxed, conversational American English: {I didn't expect to see you here.}
Dialogue language: natural Los Angeles English. The young man responds in casual local vernacular: {You really made it all the way out here?}Organize a multi-reference prompt
With a large reference set, the goal is not to squeeze every upload into one sentence. Build a small system that connects subjects, props, environments, movement, and audio.
1. Map every subject separately
Bind Character A, Character B, each product, each prop, and each location to a specific file. Avoid shortcuts such as “Images 1–4 define four people respectively.”
2. Group assets by type
Use headings such as [Characters], [Props], [Scenes], and [Motion and Audio]. Add rules that prevent identities, clothing, actions, positions, or dialogue from being swapped.
3. Build a central profile for recurring subjects
Collect appearance, fixed props, allowed locations, movement references, and exclusions in one profile when the same subject returns across scenes.
4. Choose references scene by scene
For each scene, list only the people, props, place, movement, and sound needed there. State the event and the observable end state.
[Subject Profile: Conservator]
Appearance and clothing: @Image 1.
Fixed prop: <Sample Case> from @Image 5.
Locations: <Conservation Lab> and <Gallery>.
Motion: case-opening from @Video 1; sample-placement from @Video 2.
Exclude: other characters' clothing, <Record Board>, and guide equipment.Scene 1 | Inspection in the Conservation Lab
Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the opening motion from @Video 1.
Event: <Conservator> opens the case at the workbench and inspects the sample.
End state: <Conservator> remains inside the workbench area; the case rests beside the right hand on the left side of frame.
Scene 2 | Registration in the Gallery
Use: <Registrar>, <Record Board>, and <Gallery>.
Event: <Registrar> checks the record number beside the display case.
End state: <Registrar> still holds the board with both hands; nobody else enters the display area.Multi-reference prompting is a selection problem: help the model use the right material for the current scene rather than forcing every upload to appear at once.
Plan a 30-second video in stages
For a clip with several events, divide the story into consecutive stages. Each stage should carry one main state change and end on a directly observable arrangement of people, props, and space.
[Generation Goal]
Create a <video type>. The central subject is <subject>; the overall event is <brief story summary>.
[Reference Assignments — optional]
@Image, @Video, and @Audio references define <identity, scene, motion, voice, or sound>. State what must not be inherited.
[Global Direction]
Environment and texture: <time, place, atmosphere, and physical detail>.
Visual treatment: <style, depth, palette, and lighting>.
Camera language: <framing, perspective, movement, and cutting rhythm>.
Performance focus: <the behavior or emotion that matters most>.
Exclude: <unwanted dialogue, subtitles, music, actions, objects, or distortions>.
[Stage 1]
Initial state: <people, props, and scene at the opening>.
Primary event: <one main change>.
End state: <visible positions, ownership, and scene state>.
[Stage 2]
Continue from Stage 1: <facts that remain unchanged>.
Primary event: <one main change>.
End state: <observable result>.
[Stage 3]
Primary event: <closing action>.
End state: <final visible state>.
[Maintain Consistency]
Keep <identity, character count, clothing, prop ownership, screen direction, and audio relationships> consistent.Use timestamps only when timing matters
Stages are the default. Use ranges to allocate pacing, exact time points for one critical beat, and relative timing for delays between events.
| Timing pattern | Use it like this |
|---|---|
| Time range | 0–3 seconds… 3–7 seconds… 7–12 seconds… |
| Exact time point | At 5 seconds, whip-pan left and complete the transition behind the foreground shelf. |
| Relative timing | Three seconds after the character presses the switch, the room lights fade out. |
Keep ranges consecutive and non-overlapping. They reserve time for an event; they are not frame-accurate edit points. Too little content creates extra freedom, while overloaded ranges can cause rapid cuts or skipped actions.
Scale the same structure for Long Video mode
Dreamina's dedicated Long Video mode supports a selected duration from 30 to 180 seconds. Put the chosen length and aspect ratio near the beginning, map every reference before the story, then use wider consecutive time blocks for the plot. Close with the visual, audio, and continuity rules that must remain active from start to finish.
Edit or extend an existing video
Editing, first/last-frame generation, and extension inherit some parameters from the source material. Those locked values cannot be overridden separately in the generation UI or API.
| Task | Aspect ratio | Duration |
|---|---|---|
| Video editing | Locked to the source video's ratio | Approximately preserves source duration; frame handling may create a difference of about 0.3 seconds |
| First-frame or first/last-frame | Locked to the first image; first and last images should share a ratio | Selectable |
| Video extension | Locked to the source video's ratio | Selectable for the new segment |
Editing pattern
Name one source video as the sole master. Then define the goal, target reference, edit boundary, and everything that must remain untouched.
[Edit Goal]
Edit @Video 1. Within <the full video or a time range>, <add, remove, replace, or adjust> <object, region, or audio category>.
[Source Video Role]
@Video 1 is the sole editing master. It controls characters, scene, actions, composition, camera movement, occlusion, audio, and event order.
[Target Material Role]
@Image 1 or @Audio 1 defines only <the target attributes>.
[Edit Scope]
Change only <object, region, time range, or sound category>.
[Preserve]
Keep <all visuals, movement, audio, and timing relationships that must not change> from @Video 1.When a box, brush, arrow, or point marks the edit area, mention the annotation but still name the target. For example: “Inside the marked area, replace the boy's blue jeans with tailored black suit trousers during the first 8 seconds.” The mark locates the region; the text defines the object, change, and timing.
The same structure works for subject replacement, background replacement, lighting changes, dialogue language, music removal, and sound-effect edits. For an object swap, state the exact object count and require the replacement to inherit the original object's motion, occlusion, entry, exit, path, and timing.
Forward and backward extension
A forward extension begins from the source's last frame; a backward extension must end on the source's first frame. Describe the shared boundary state before introducing any new event.
@Video 1 is the source to extend forward.
Continue directly from its last frame. Preserve subject pose and orientation, prop positions, background geometry, camera composition, lighting, and motion direction.
Then <describe the new action, camera treatment, and audio>.
Keep identities, clothing, key props, screen direction, and object count continuous. Do not duplicate or split a subject.Use keyframes, storyboards, and visible direction
Build a recognizable, believable character
Avoid relying on a generic label such as “cinematic woman” or “handsome man.” Establish identity with a small set of stable, observable traits. For realistic people, combine age and visual identity, natural skin detail, 3–4 facial landmarks, gaze, hair, clothing, body shape, posture, and temperament. The same framework also works for stylized or animated characters.
Identity: <age, cultural or ethnic appearance, and overall facial character>.
Skin: <undertone, complexion, pores, freckles, lines, or other natural texture>.
Face: <3–4 distinguishing features across eyes, brows, nose, lips, cheekbones, or jaw>.
Gaze and emotion: <where the character looks, how focused the eyes are, and the feeling conveyed>.
Hair: <color, length, style, texture, condition, and interaction with wind or movement>.
Wardrobe: <cut, color, garment, fabric, condition, and defining details>.
Presence: <build, posture, framing, action, and overall temperament>.Physical cues are more useful than abstract emotion alone. Instead of only writing “heartbroken,” direct the lowered gaze, restrained breath, tightening jaw, trembling lips, delayed response, or other behavior that makes the feeling visible on screen.
Describe how a transition actually happens
A transition name is only a starting point. Define the method, continuity safeguards, the visual bridge out of Shot A, and the composition that should be visible when Shot B arrives.
| Transition | Prompt the visible mechanism |
|---|---|
| Natural cut | Let Shot A complete its readable action, then move into a related action in Shot B without an abrupt jump or a newly appearing object. |
| Fade | Allow Shot A to darken or brighten fully before Shot B emerges, commonly to open, close, or signal elapsed time. |
| Occlusion or mask | Move a foreground object across the lens until it fills the frame; reveal the new scene as that cover passes or the camera pulls away. |
| Shape or color match | End Shot A on a dominant shape or color and begin Shot B with a visually corresponding form in the same area of the frame. |
| Whip-pan or motion bridge | Carry the same movement direction and energy through full-frame motion blur, then settle into the destination shot. |
First and last frames
Assign @Image 1 as the first frame and @Image 2 as the last frame in separate statements. Use matching aspect ratios. Other images may supplement identity, props, materials, or scene details without replacing the anchor compositions.
Multiple keyframes
State that @Image 1 through @Image N are keyframes in order, then define the important state in each image. Separate images are generally easier to align than a collage.
Storyboard grids
Explain the panel order, shot structure, action, and continuity to inherit. A grid is guidance for composition and sequence, not an instruction to reproduce every panel as an exact frame.
Blockout references
Identify whether the blockout is coarse or fine. State which timing, movement, camera, spatial structure, materials, and style should carry into the render.
One-click video and transitions
Define image order, material roles, motion amount, editing rhythm, and audio. For a transition, describe the trigger, occluding object, camera direction, transition mechanism, and arrival state.
Emotion and performance
Translate abstract emotion into visible behavior: gaze, breath, posture, facial tension, hesitation, gesture, timing, and the event that causes the change.
Cinematography terms
Pair terms such as tracking shot, shallow depth of field, golden hour, natural vignette, or whip-pan with the visible outcome you want. Numeric lens or shutter values alone are less clear.
@Image 1 is the first frame. It controls the opening composition, subject position, pose, prop state, scene, and camera direction.
@Image 2 is the last frame. It controls the ending composition, subject position, pose, prop state, scene, and camera direction.
@Image 3 defines <Subject A>'s appearance or clothing without changing either anchor composition.
Describe one continuous event that begins naturally at @Image 1 and resolves at @Image 2. Maintain identity, prop structure and ownership, scene layout, and camera direction between the anchors.Check the prompt before generating
- The subject and primary event are explicit.
- Every reference says what to use and what to ignore.
- Each character, product, and prop has a name and file mapping.
- References are selected per scene rather than forced into every scene.
- Each long-video stage has one main change and a visible end state.
- Identity, clothing, object ownership, count, and spatial direction stay consistent.
- An editing prompt names one master, a narrow edit scope, target quantity, and preserve list.
- Emotion and camera terminology are paired with visible or audible cues.
- First/last frames and keyframes each have one role; anchor images share an aspect ratio.
- Storyboard or blockout prompts state the exact structure and detail level to inherit.
- Locked aspect-ratio and duration rules match the chosen task.
- Extensions check boundary image, motion trend, object continuity, and audio.
Important limitations
- Timestamps budget time; they do not define frame-accurate edit points.
- Editing prompts can improve event alignment but cannot promise frame-by-frame overlap.
- Multi-reference creation selects and combines useful assets; it does not guarantee that every upload appears.
- Exact subtitles, formulas, signs, product specifications, and frame-level timing may still require prepared assets and post-production.
- First/last-frame generation may stretch the last image when the two anchor ratios differ.
- Seamless transitions target perceptual continuity, not pixel-identical preservation of both source clips.
Results vary with source quality, task complexity, selected settings, and the relationships among reference materials. This guide explains prompt-writing technique rather than guaranteeing a particular output.
