Nothing in this clip was filmed. No camera, no studio, no shipping a product to a creator and waiting nine days. One production run, thirty seconds, $7.76.
We took a competitor's viral ad, kept its format, and replaced the person with our presenter and the product with ours. Then we asked five video models for the same scene and paid for every one of them.
Below is what came back, what each one broke, and what the finished thing costs.
The job is harder than "generate a video"
Two things have to survive every single frame: a face and a label.
Miss the face and it stops being your presenter. Miss the label and it stops being your product — which is the entire reason the ad exists. Every model here can produce a pretty fifteen seconds. Almost none of them can produce your fifteen seconds.
That is the test. Not motion quality, not resolution, not what the demo reel on the vendor's homepage looks like.
Five models, one starting frame, no cherry-picking
Same first frame, same instruction, first result kept in every case. Prices are what you would pay to run each clip on OmniAgent.
What each one actually did, and where it broke
Models do not fail generically. Each has one specific limit, and that limit is what picks the model for you.
Seedance 2.0 — $0.25/second — the one we ship
It takes an ordered list of references. Starting frame, the presenter's
identity card, the product card — addressed in the prompt as @Image1,
@Image2, @Image3. That is why it holds the face and the packaging together:
it is not guessing which reference is which.
Nothing else in this test can do that, and everything else in this test loses something because of it.
Its limits, all of which cost us money to find:
- It needs a legend, or it improvises. Without a sentence naming what each reference is, the live run painted the product's colours onto the presenter's shirt. Right images, wrong idea of them.
- The starting frame is not a reference. It travels in the provider's own first-frame slot. Flatten the two ideas together and you drop a frame — we did, on a live reel.
- A video input re-prices the call. Feed it footage and the rate applies to input plus output seconds — roughly 10% more per attempt. Sometimes that is worth paying (below); it is just never free.
- It costs five times the cheap lane. That is why the cheap lane exists.
Grok Imagine 1.5 — $0.05/second — the draft lane
85–90% of the Seedance result for a fifth of the price. At 1080p it accepts exactly one reference, which sounds disqualifying and mostly is not: the face and the product are already baked into the starting frame, so a second reference has nothing left to carry.
Its limits:
- One reference means no path to correct anything. If the starting frame got the product slightly wrong, nothing downstream fixes it. With Seedance you add a product card and it recovers.
- The missing 10–15% is exactly what people notice — label edges, the geometry of a hand on a wrapper.
Fine while you are still deciding what to make. Not fine for the cut that ships.
PixVerse V6 — $1.66 for the clip
Packaging came out fine — competitive with Seedance on label fidelity.
What ruled it out: the hero moment came out weak. In this ad the hero moment is the caramel stretch, the single second the whole thing exists for. A model that renders everything competently and the payoff limply is a model for filler shots, not for the shot.
MiniMax H3 — $2.02 for the clip
The best caramel stretch of anything we ran. Texture and physics, clearly ahead of the field.
What ruled it out: the worst fidelity to the packaging. It reinvents label details that must not move. If your hero moment is a texture and your product is unbranded, this comparison ends differently and H3 wins it. Ours is a branded wrapper with exact typography, so it lost on the one axis we could not compromise.
Kling motion-control — $0.66 for the clip
It answers a different question: it transfers movement from a reference clip instead of inventing it from a prompt.
Its limits: it returns no audio at all, and it needs a source video to copy motion from — so on a replica brief the only clip available is the competitor's, with everything that drags in. Real tool, wrong job here. Keep it for when you have a movement of your own that must be matched exactly.
The pattern worth stealing
Every model was best at something. The one we ship is best at nothing.
H3 beats it on texture. Grok beats it on price. PixVerse matches it on packaging. Seedance wins because it is the only one that does not fail on the axis we cannot compromise — and because multi-reference input is the difference between correcting a mistake and re-rolling the dice.
Pick on your own non-negotiable axis. Not on an aggregate score, and definitely not on a demo reel.
What a finished part actually costs
Per 15-second part, measured on the same scene, at OmniAgent prices:
| Item | Ship it (Seedance 2) | Draft it (Grok 1080p) |
|---|---|---|
| Starting frame (nano-banana) | $0.11 | $0.11 |
| Video, 15s | $3.69 | $0.72 |
| Voice (ElevenLabs STS) | $0.08 | $0.08 |
| Total per part | $3.88 | $0.91 |
The gap is 4.3×, and the decision follows from it.
A finished piece someone will actually publish is worth $3.88. Twelve attempts at working out the concept are not — that is about $11 on Grok against $47 on Seedance, for twelve versions of a thing you are going to throw away eleven of.
These are the 5 August prices. The family has since been re-priced and grew a new generation — what the August catalog drop changed is its own measured story, promo deadlines included.
Three things that cost money to learn
Handing the model the source video works — it just costs you twice. Showing it the original is the obvious way to copy a format, and it does carry the motion faithfully. What it also carries is the other brand: across our first three attempts the competitor's packaging bled into our render every time, and each one took another pass of prompt work to push back out. On Seedance a video input additionally re-rates the call as input plus output seconds — about 10% more per attempt, which compounds with the extra attempts.
So it is a tool, not a trap: reach for it when the movement is hard to describe and you are willing to iterate. Our default is the other way round — the motion goes into the prompt, beat by beat — because that starts clean instead of starting with something to remove.
Regenerate the broken seconds, not the clip. When a defect lands at second
eleven the instinct is to re-run everything. Find the speech pause with ffmpeg silencedetect, cut there, generate only the tail using that frame as the new
start, and splice. On the live run: $1.73 instead of $3.88.
Let the model speak, then swap the voice. Both Seedance and Grok generate speech with native lip-sync — put the line in the prompt, in quotes. Convert to the presenter's cloned voice afterwards with speech-to-speech, which preserves timing to within 1–6 milliseconds, so the sync survives. Video-to-video lip-sync models are a step you do not need.
A smaller trap worth writing down: generated clips come back with an audio track about 19 ms shorter than the video. Concatenate a few and the sound walks away from the picture. Pad each piece to its exact video length before joining.
Where this sits in an AI content automation pipeline
A model comparison is only useful if something runs it on a schedule. The clip above is one step of five, and the other four are where the time goes:
- Find the format. A competitor scan reads what is landing in your niche right now and turns it into a brief — this ad came from one. Reading a reel frame by frame is free, and what a teardown actually measures is thirty cuts made out of two clips.
- Build the identity. A presenter card and a product card, generated once from real photos, then reused by every run so the face and the packaging stop being a per-clip gamble. Skip it and the model invents your brand — the failure is measured in what it draws when it has never seen your product.
- Price the run before it starts. Every model, every stage, in dollars, before anything is charged. This is the step most tools skip, and it is the one that keeps a month of AI content automation from quietly becoming a four-figure invoice.
- Generate, gate, regenerate. Score the hook before spending on assets, and re-run only the seconds that broke.
- Publish on a schedule. The finished part goes out to Instagram, TikTok and YouTube without a human moving a file.
The model choice above decides step 4. Steps 1, 2, 3 and 5 decide whether you publish twice a month or twice a day. And the machinery that connects the five — which agent checks which output, and where a human approval actually sits — is its own discipline: what 134 production bug reports taught us about agent orchestration.
What we would tell someone choosing today
- Run two models, not one. A cheap lane for deciding what to make, an expensive one for making it. This single habit is worth more than picking the "best" model.
- Choose on your failure mode, not on a demo. Ours is a branded wrapper, so packaging fidelity decided it. If your hero moment is a texture, H3 wins the same comparison.
- Price the part, not the second. The frame and the voice add $0.19 whichever video model you pick — worth knowing before you optimise the wrong number.
- Iterate on frames, not clips. A starting frame costs $0.11. Getting it right before spending $3.69 on motion is the cheapest habit in the pipeline.
This is the format our autopilot runs on a schedule, priced before each run rather than after — see what a run costs or watch the agent plan one.

