OmniAgentby omnigems.ai

Four video models, one shot: what a real ad replica cost us

We rebuilt someone else's viral ad with our own presenter and product, then ran the same scene through Seedance 2, Grok Imagine, PixVerse and MiniMax H3. Measured prices, measured failures, and the one we actually ship.

Short version: Seedance 2.0 is what we ship, Grok Imagine is what we experiment with. Seedance holds a face and a product label better than anything else we tested, at about $0.205 per second. Grok reaches roughly 85–90% of that for a fifth of the price, which makes it the right tool for iterating and the wrong one for the final cut.

Everything below comes from one production run: a competitor's viral ad, rebuilt with our own presenter and our own product. Here is what came out of it.

What we shippedSomeone else's ad, our presenter, our productTwo 15-second parts on Seedance 2.0, the presenter's own cloned voice over the model's native lip-sync. Nothing here was filmed. Click it to watch with sound.$7.76 total30 seconds1 production run

The task

Take an existing ad. Keep the format — framing, pacing, the beat the whole thing exists for. Replace the person with our presenter and the product with ours, and make it indistinguishable from something we filmed.

That is harder than "generate a video", because two things have to survive every frame: a face and a label. Models fail those two differently, and the differences are the whole comparison.

The four, side by side

Same starting frame, same instruction, no cherry-picking. Prices are our catalogue rate multiplied by each clip's measured length.

grok-imagineimage-to-video · 1080p · 15s$0.72
kling 2.6motion-control · 720p · 10s$0.66
MiniMax H3image-to-video · 768p · 15s$2.02
pixverse v6image-to-video · 1080p · 15s$1.66
One starting frame, one instruction, four models. Prices are what you would pay to run each clip here — our catalogue rate × the clip's measured length. Click any of them to watch full size with sound.

What each one actually did, and where it broke

The models do not fail generically. Each has a specific limit, and the limit is what makes the choice for you.

Seedance 2.0 — $0.205/s — the one we ship

What it does that nothing else here does: it accepts an ordered list of references. Starting frame, the presenter's identity card, the product card — addressed in the prompt as @Image1, @Image2, @Image3. That is why it holds both the face and the packaging: it is not guessing which reference is which.

Its limits, all of which cost us something:

  • It needs a legend, or it improvises. Without a sentence naming what each reference is, the live run watched it paint the product's colours onto the presenter's shirt. It had the right images and the wrong idea of them.
  • The starting frame is not a reference. It travels in the provider's own first-frame slot. Flatten the two concepts together and you drop a frame — we did, on a live reel.
  • A video input re-prices the call. Feed it footage and the rate applies to input plus output seconds, roughly 10% more for a thing you should not be doing anyway (below).
  • It is five times the price of the cheap lane. That is the entire reason the cheap lane exists.

Grok Imagine 1.5 — $0.04/s — the experiment lane

Roughly 85–90% of the Seedance result for a fifth of the price. At 1080p it accepts exactly one reference, which sounds disqualifying and mostly is not: the face and the product are already baked into the starting frame, so a second reference has nothing left to carry.

Its limits:

  • One reference means no correction. If the starting frame got the product slightly wrong, nothing downstream can fix it — with Seedance you add a product card and it recovers.
  • The last 10–15% is exactly the part people notice — label edges, the precise geometry of a hand on a wrapper. Fine while you are deciding what to make. Not fine for the cut that ships.

While you are still working out what the video should even be, paying five times more per attempt is the actual mistake.

PixVerse V6 — $1.38 for the clip

Packaging came out fine — genuinely competitive with Seedance on label fidelity.

The limit that ruled it out: the hero moment came out weak. In this ad the hero moment is the caramel stretch, the single second the whole thing exists for. A model that renders everything competently and the payoff limply is a model you can use for filler shots and not for the shot.

MiniMax H3 — $1.69 for the clip

The best caramel stretch of anything we ran. Texture and physics, clearly ahead of the field.

The limit: the worst fidelity to the packaging. It reinvents label details that must not move. If your hero moment is a texture and your product is unbranded or generic, this comparison ends differently and H3 wins it. Ours is a branded wrapper with exact typography, so it lost on the one axis we could not compromise.

Kling motion-control — $0.55 for the clip

Answers a different question entirely: it transfers movement from a reference clip rather than inventing it from a prompt.

The limits: it produces no audio at all, and it needs a source video to copy motion from — which, for replicating someone else's ad, is precisely the input we refuse to use. Real tool, wrong job here. Keep it for when you have a movement of your own that must be matched exactly.

The pattern

Every model was best at something. The one we ship is not the best at any single axis — H3 beats it on texture, Grok beats it on price, PixVerse matches it on packaging. Seedance wins because it is the only one that does not fail on the axis we cannot compromise, and because multi-reference input is the difference between correcting a mistake and re-rolling the dice.

Pick on your own non-negotiable axis, not on an aggregate score.

What a finished part actually costs

Per 15-second part, measured on the same scene:

ItemHero (Seedance 2)Cheap (Grok 1080p)
Starting frame (nano-banana)~$0.10~$0.10
Video, 15s~$3.08$0.60
Voice (ElevenLabs STS)~$0.07~$0.07
Total per part~$3.25~$0.77

The gap is 4×, and it is the whole decision. A finished piece someone will publish is worth $3.25. Twelve attempts at working out the concept are not — that is $9.24 on Grok against $39 on Seedance.

Three things that cost money to learn

Never hand the model the source video. It seems obvious that the fastest way to copy a format is to show the original. We tried three times. The competitor's brand leaked into our render every time — and on Seedance a video input also re-rates the call as (input + output) seconds, so the leak costs about 10% extra for the privilege. The motion goes into the prompt instead, beat by beat.

Regenerate the broken seconds, not the clip. When a defect lands at second eleven the instinct is to re-run everything. Find the speech pause with ffmpeg silencedetect, cut there, generate only the tail using that frame as the new start, and splice. On the live run: $1.44 instead of $3.25.

Let the model speak, then swap the voice. Both Seedance and Grok generate speech with native lip-sync — put the line in the prompt, in quotes. Convert to the presenter's cloned voice afterwards with speech-to-speech, which preserves timing to within 1–6 milliseconds, so the sync survives. Video-to-video lip-sync models are a step you do not need.

A smaller trap worth writing down: generated clips come back with an audio track about 19 ms shorter than the video. Concatenate a few and the sound walks away from the picture. Pad each piece to its exact video length before joining.

What we would tell someone choosing today

  1. Pick two models, not one. A cheap lane for deciding what to make, an expensive lane for making it. Running everything through the good model is how a $40 experiment becomes a $400 one.
  2. Choose on your failure mode, not on a demo reel. Ours is a branded wrapper, so packaging fidelity decided it. If your hero moment is a texture, H3 wins the same comparison.
  3. Price the part, not the second. The frame and the voice add about $0.17 whichever video model you pick — worth knowing before optimising the wrong number.
  4. Iterate on frames, not clips. A starting frame costs about $0.09. Getting it right before spending $3 on motion is the cheapest habit in the pipeline.

This is the format our autopilot runs on a schedule, priced before each run rather than after — see what a run costs or watch the agent plan one.

Questions

Which AI video model is best for product ads?
In our testing, Seedance 2.0 — the only one that accepts an ordered list of
How much does one AI-generated ad clip cost?
A finished 15-second part cost us about **$3.25** on Seedance 2.0 and about
Can you use the original video as a reference to copy a format?
We do not recommend it. Across three attempts the source brand leaked into the
How do you fix a defect without paying for the whole clip again?
Find a speech pause with `ffmpeg silencedetect`, cut there, and generate only the
Do these models handle speech and lip-sync themselves?
Yes — both Seedance and Grok generate the line with native lip-sync when it is

Omniagent / ai content factory

Run this pipeline on your own product

Every model in this comparison is one dropdown in the same system — with the price shown before the run, and the cost logged after it. Start on the free plan and pay only for what you generate.

Describe the productApprove the framesPublish on a schedule
Four video models, one shot: what a real ad replica cost us