Eighty men in identical charcoal hoodies stand in rows on wet concrete steps. On the beat they fold forward, all of them, and snap back up. One man in a white t-shirt does not. He lights a cigarette.
Nobody choreographed that: the movement, the slow pull-back and the timing of the bows come from a 30-second reference clip. The text described only who is standing there, what they wear, where they are and how the light falls. This is the cheapest kind of video to make well — motion is the hard part, and you are not making it, you are borrowing it.
Two ways below. The short one: two steps in the OmniAgent chat with your own avatar. The long one: the same prompt and settings, if you want to build it by hand in any other generator.
The STORM format — one continuous shot0:30 · 16:9
Borrow the motion, write the rest. The avatar's card is made once for about 12 credits; after that a reference-video shot is one message and an approved price. Start in the chat — or see what a run costs first.
Two steps on OmniAgent
Step 1 — the avatar
The shot has one lead: your avatar from the roster. If there is none yet, start with the identity card: four free configurator steps and one generation of about 12 credits, after which the avatar has a face that reads from any angle. The card is made once; every format after it simply uses it.
The start frame — the lead chest-up, white t-shirt, cigarette in the lips, a soft crowd behind — is not yours to draw: the agent makes it from the card for 50 credits, and the shot opens on that composition.
Step 1 — the avatar's identity card

Full body from four sides on top, the face from the same four below. Made once, about 12 credits; every format after it reads this card as the face.
Step 2 — the start frame the agent draws

Drawn from the card by the agent, 50 credits: white t-shirt, cigarette in the lips, a soft crowd behind. The shot opens on this composition — you do not make it yourself.
Step 2 — one message
Attach the reference clip (where to get it is in the do-it-yourself section) and tell the agent in one message what you want. Copy this as it is, with your avatar's name in it:
Make the STORM format with the avatar <avatar name>: Seedance 2.5, 30
seconds, 16:9, 480p, no audio. The attached video is the reference for
movement, camera and timing; everything else comes from the prompt. Draw
the start frame from the card: white t-shirt, cigarette in the lips, a crowd
in charcoal hoodies on concrete steps.
The agent builds the spec, shows the quote — 638 credits for 30 seconds at 480p — and runs nothing until you say go. A few minutes later the shot is in Content next to everything else, and from there it goes to review, the schedule and publishing.
Want your own crowd, your own place or a different colour on the lead — add it to the same message. What to change and what to keep is in the last section.
What it costs
Prices are the platform's credits, as the quote shows them before the run. One credit is one cent.
| What | When | Credits | ≈ USD |
|---|---|---|---|
| Avatar identity card | once, reused after | 12 | $0.12 |
| Start frame | once per format | 50 | $0.50 |
| The shot, 30 s, Seedance 2.5, 480p | per take | 638 | $6.38 |
| The shot, 30 s, Seedance 2.5, 720p | per take | 1 181 | $11.81 |
Why 30 seconds cost that much: the model bills the reference seconds too — 30 in plus 30 out. A long reference raises the bill rather than cutting it, so the clip is trimmed to the seconds you actually want to borrow. The full Seedance price sheet and the rest of the lineup are in the cheapest AI video generator piece.
Do it yourself in any generator
If you run your own generator with reference-video support — Seedance 2.5 through the API, or any service that carries it — you need three things: the clip, one image of the lead, and the prompt below.
Where the reference clip comes from
The clip is 30 seconds of GENER8ION — STORM with Yung Lean: a packed crowd doing one sharp repeated movement, one still person in the middle. Cut the section with the bows and bring it to three rules: 30 seconds or less, no audio track, under about 10 MB — a free online compressor with "720p" and "remove audio" ticked does all three. Do not use a fan edit from a feed: burned-in subtitles and watermarks are read by the model as part of the reference.
Any other clip works as long as the movement is clear. Good references have a lot of motion and few ideas — a crowd doing one thing, a person walking toward the camera, a head turning. Bad ones cut a lot and never repeat a shape.
The lead's image
One sharp, front-facing chest-up frame of the lead, already in the right clothes — the shot opens on it. On OmniAgent the agent draws it from the card; elsewhere, make it with any image model using a photo of the face as the reference. One image is enough; do not add more.
The prompt
Change the lead, the crowd's colour and the place; keep the shape of the lines.
Keep the movement, camera, pacing, timing and cut points from the reference
video. Replace everything else.
The attached image is the lead. He is a young man, dark curly hair,
moustache, plain white t-shirt #f2f2f0, dark jeans, a cigarette in his
lips. Keep his face, bone structure and hair exactly as shown — do not
restyle him. He stands dead centre, arms at his sides, eyes straight into
the lens. He is the only one in white and the only one smoking. Only ONE
of him at any time.
THE CROWD is everyone else — about eighty tech-conference people, packed
edge to edge, filling every row front to back. Men and women mixed, ages
25 to 60, real variation in face, height, age and build, so the same face
never sits next to itself.
Every one of them wears the SAME outfit: a plain charcoal #2b2f33 zip-up
hoodie, same cut, zipped closed, over a black crew-neck t-shirt, black
trousers, black low shoes, and the same white lanyard with a blank white
badge hanging on every chest. The cloth colour does not change: all of
them in flat charcoal.
No ties. No neckties of any kind, no striped ties, no knitted ties. No
school blazers, no suit jackets, no shirt-and-tie, no uniformed
schoolboys. Nobody else smokes.
THE PLACE is an open-air amphitheatre of raw grey concrete — broad shallow
steps rising back and up, wet from rain, and one blank poured-concrete
wall behind the top row. Overcast light, no sky in frame.
NO brick. No red brickwork, no arched doorway, no school building, no
barred windows, no institutional facade. The wall behind the top row is
blank concrete and nothing else. No signs, no text, no subtitles, no
captions, no watermark, no logos anywhere in frame.
THE ACTION is exactly as the reference plays it: the crowd convulses in
unison, sharp jerks of head and shoulders, folding forward and snapping
upright on the beat. The lead does not convulse, does not fold, does not
bow. He does what the still figure in the reference does — he brings a
lighter up to his face, lights the cigarette in his lips, draws on it,
lowers the hand and lets the smoke go. That is his only movement.
LOOK — full colour, cold, dark.
One overcast source from directly overhead, falling off fast — lit
shoulders, shaded faces, near black in the gaps between bodies.
Lead and highlight: #f2f2f0 #ff5a00
Cold and drained: #e8ebee #9aa3ab #4a5159 #2b2f33 #1c2127
Skin: lit #e8c3a8, shadow #6b4a3c
Very dark overall. The brightest point is the lit shoulder line of the
lead's white t-shirt, and it never reaches pure white. Clipped white must
never appear. The cigarette ember is the only warm point in the frame.
No navy, no golden sunlight, no teal-and-orange, no magenta, no purple
anywhere.
Deep shadow, soft rolled highlights. No glow, no haze, no bloom, no grain.
FRAME — 16:9, the crowd filling the frame edge to edge, rows receding
back.
The settings
For Seedance 2.5 through the API: the lead's image as the first reference, the clip as the reference video, and no separate first-frame field:
{
"model": "bytedance/seedance-2-5",
"input": {
"prompt": "<the prompt above>",
"reference_image_urls": ["<URL of the lead's image>"],
"reference_video_urls": ["<URL of the reference clip>"],
"generate_audio": false,
"resolution": "480p",
"aspect_ratio": "16:9",
"duration": 30,
"output_format": "mp4"
}
}
Look at the first result at 480p — it shows the movement, the pull-back and the one person who does not bow; judge the face and the crowd at 720p with the same request.
Five things that do the work
One sentence hands over the motion. Movement, camera, pacing, cut points — named as a group in the first line, given to the video, never mentioned again. The moment the text describes a dolly or a beat, the shot has two authors, and the reference takes not just the camera but its own wardrobe.
Hex codes, not colour names. #2b2f33 is one charcoal; "dark grey" is a
different grey every run. Every colour that matters is a code, including the
skin in light and in shadow.
Counts. "About eighty" fills a 16:9 frame front to back. "Only ONE of him" is there because the lead is the most described person in the prompt, and a model likes to make more of what it is told most about.
Refuse things twice. The reference has its own clothes and its own building. Whatever it keeps dragging into the frame is refused in two different sentences — and stays out to the last second.
Name the brightest point. "The lit shoulder line of the lead's white t-shirt, and it never reaches pure white." Without it the whole image lifts towards grey to make room for the white.
A seedance prompt by reference: what to make your own
Three lines carry the whole look — change them and the same reference gives a different video.
- The lead's colour.
#f2f2f0and "the only one in white" — swap for your own; put the same hex in the start-frame description, or the lead will not match himself. - The crowd's colour. Charcoal here. Any flat colour works as long as it is far from the lead's: the contrast between one and everyone is the only axis the shot stands on.
- The place. Materials and shapes only — concrete, steps, wet, a blank wall. Then the original location's most stubborn feature, named and refused twice.
Keep the counts and the structure: "about eighty", "only ONE of him" and the garments listed part by part are doing the same job whatever the subject is. On OmniAgent all of that is the same lines in one message to the agent; the quote shows the price before the run, and the shot lands in Content. Plans and included credits are on the pricing page.

