Get the full report

Enter your details to unlock the full guide.

Start free
For growth and creative teams

AI Video Prompts That Sell,
Not Just Look Good

Craft is the floor. Direction is the job.

Turn raw AI video into ad creative that converts, not just pretty clips.
Start with the five failures to fix, then the method, the current models, and a free prompt builder.

14 min read · By Quickads

The 5 usual suspects

5 ways your AI video goes wrong

Most guides teach the method first and hope you avoid the potholes. This one opens with the potholes. Below are the five ways AI video prompts actually fail, each named the way it shows up on your screen. Open the one that matches what you are looking at, then read the rest of the guide as the how-to for avoiding all five. Each includes what you are seeing, why it happens, and the exact fix with a fragment you can paste straight in.

01It looks generic

Your clip looks like everyone else's because you described a scene instead of directing a shot. The fix lives in the cinematography and style parts of your prompt: name the camera, the lens, and the light.

What you're seeing

The output is competent and completely forgettable. Soft, even lighting, a slow generic drift, a subject that could belong to any brand. It looks like AI video because it looks like the average of all AI video.

Why it happens

You wrote a description, not a shot. The two things that set a look, the camera and the light, are missing. When you leave those out, the model fills the gap with its default, and its default is the middle of everything it has ever seen. Words like "cinematic" and "moody" make it worse, because they sound like direction while telling the model nothing it can act on.

The fix

Name the shot and the light in concrete terms. Pick a camera move, a lens, and a specific quality of light, then let those choices carry the look. Trade every vague adjective for a real one a cinematographer would recognize.

Copy-paste fragmentSlow dolly-in on [product], [one specific action]. [Setting], warm side light, 50mm lens, shallow depth of field, subtle 35mm grain. 9:16 vertical.
02The product changes between shots

The product drifts between shots because you worked text only when the look had to stay locked. Start from a reference image so the packaging, color, and shape hold before the model animates a frame.

What you're seeing

The label morphs, the cap changes shape, the color shifts between one generation and the next. Across a set of variations, it stops looking like the same product, which for a brand is the whole problem.

Why it happens

This one is not a missing prompt part, it is the wrong method. You generated from text only, so the model reinvents the product every single frame because it has nothing fixed to hold on to. No amount of extra description will pin down a shape the way a reference image does. This is the difference between generating from text alone and generating from a locked reference image, and it is why a reference is not optional for product work.

The fix

Start from a locked reference image of the real product, then state plainly what must stay identical before you describe any motion. Lock first, move second.

Copy-paste fragmentUse this reference image. Keep the bottle shape, matte-black finish, and label exactly as shown. Then: slow push-in as the dropper lifts. Do not restyle or redraw the packaging.
03The motion falls apart

Limbs bend, logos warp, the scene melts. Motion breaks when the action you asked for is too big or self contradictory. Keep it small, motivated, and free of conflicting instructions.

What you're seeing

Hands grow extra fingers, the product bends like rubber, the background swims, or a fast camera move smears the whole frame. It looks fine on the first frame and dissolves by the last.

Why it happens

The break is in the action you asked for. Big, fast, complex movement is still where models fall apart, so an over-ambitious action gives the model more than it can keep coherent. Conflicting instructions do the same damage: "keep the product identical" and "make it transform" cannot both be true, and neither can "no movement" and "dramatic action". The model tries to satisfy both and warps.

The fix

Ask for one small, motivated motion and lock the camera unless you have a reason to move it. Remove any instruction that fights another. Small motion holds up on screen, and it reads as more premium than a chaotic one anyway.

Copy-paste fragmentMedium shot, locked camera. [Subject] with one small motion only: the cap lifts and the dropper rises two inches. No transformation, no fast pans, no scene changes.
04It's beautiful and it doesn't sell

The footage is beautiful and it does nothing. Every craft part was right but the ad layer was missing: no first-second hook, no product truth, no single reason to act.

What you're seeing

A genuinely gorgeous clip that performs like a screensaver. People watch, maybe, and then nothing happens. No clicks, no saves, no lift on the number you actually care about.

Why it happens

All the craft was there, so the shot is clean. What is missing is the ad layer on top: a hook, the product truth, and one clear action. There is no hook in the first second, so the scroll never stops. The product is never shown doing its real job, so there is no reason to believe it. And there is no single clear action, so even an interested viewer has nowhere to go. Beautiful is the floor now, not the goal.

The fix

Point the shot at an outcome. Open on motion or the product mid-frame in the first second, show the product doing the one thing it is for, and end on a single, readable call to action. One ask, not three.

Copy-paste fragmentOpen on [product] mid-frame in the first second. Show [the product doing its real job] by second two. On-screen text "one clear line" from second 2 to second 4. One call to action, nothing else.
05You burned credits on the wrong output

You spent credits and got the wrong file back. The prompt was fine; the setup was not. Match the aspect ratio to your reference and name the model before you generate.

What you're seeing

The clip is good, but it is the wrong shape for where it has to run, or the phrasing was tuned for a model you were not actually using. Either way you are paying to generate again, and video credits are not cheap.

Why it happens

This is not a prompt-part failure, it is a setup failure that happens before the model even reads your prompt. A reference image in one aspect ratio and a generation in another wastes the credit outright. And when an LLM writes the prompt without being told which tool it is for, the phrasing lands slightly wrong, because the best wording genuinely differs by model. Small setup slips, expensive on repeat.

The fix

Run a ten-second setup check before every generation. Match the aspect ratio to your reference, and tell the tool, or the LLM writing for it, exactly which model the prompt is for.

Copy-paste fragmentModel: Seedance 2.0. Reference image: 9:16. Generation: 9:16 vertical, matched to the reference. Confirm the aspect ratio and model before rendering.

You are describing. Start directing.

Type "a woman drinking coffee in a sunny kitchen, cinematic" into any video model and you will get something back. It will also look generic, because you described a scene and left every real decision to the model. An ad director does not do that. An ad director decides the hook timing, the pattern interrupt, the product moment, and the exact frame where the viewer stops scrolling.

That gap is the whole game. Movement, pacing, and camera behavior are the details that separate a clip that looks like everyone else's from one that looks like yours.

For an ad there is a second layer. A pretty clip is not the job. The job is to stop the scroll in the first second, show the product doing something real, and give one clear reason to act. You can generate beautiful footage all day. If it is not pointed at that outcome, it is all gas and no steering.

This guide covers both: the craft of directing AI video, and the ad layer that makes the footage sell. It is grounded in what Quickads sees across 30,000+ brands and an ad library of 30M+ ads, where the same pattern holds. The videos that perform are directed, not described.

Text-to-video vs. image-to-video

There are two ways to generate AI video, and they are not equal.

Text-to-video hands the model the wheel. You give words, it invents the rest. That is fine for loose exploration, but for anything that has to look a specific way it burns credits on outputs that miss.

Image-to-video is the one to learn. You start from a reference image, usually one you generated and locked first, so the look is set before the model animates a single frame. For product and ad work this is not optional. It is how the product stays on-brand and recognizable across every variation. It is also why image-to-video is the backbone of the creative work inside Quickads: your real product packaging stays consistent while the scene, the motion, and the hook change across dozens of variations.

The method has moved on since most guides were written. Reference control is no longer one image. The strongest models today take several references at once, images, short clips, even audio, and that, not a longer prompt, is how you hold a character or a product consistent across shots.

One rule for image-to-video: say what must stay the same before you describe any motion. Lock the product shape, color, and packaging, then add small, motivated movement. Small motion holds up. Big dramatic motion falls apart.

Pro tip. Match your aspect ratio to your reference image before you generate. Video credits are expensive, and the wrong format is wasted spend.

Pick your model

There is no single best AI video model. There is a best model for the job in front of you. Here is the current picture as of July 2026. This space changes monthly, so treat it as a snapshot, not a law.

ModelBest forWhat to know
Seedance 2.0The default E-commerce product drops and everyday ad work Strong prompt adherence, believable physics and motion, native audio, and clips up to around fifteen seconds. Start here.
Google Veo 3.1Realism Founder and UGC talking-head ads Best-in-class realism and the standout for synchronized spoken dialogue, not just sound effects. Reach for it when the ad needs a voice.
Kling 3.0Stylized Stylized, cinematic brand spots Native 4K, multi-shot sequences, and multilingual lip-sync. The pick when you want a distinctive, art-directed feel.
Runway Gen-4.5Control Precise, art-directed hero shots Motion brushes and explicit camera choreography, plus a video-to-video mode to edit shots you already have.
Adobe Firefly VideoCompliance Regulated industries and legal safety Trained on licensed content, so it is the conservative choice when the legal ground has to be solid.

One more, so you are not working from an old list: OpenAI has wound Sora down. The app is closed and API access ends in September 2026, and most commercial work has moved to Seedance 2.0 and Veo 3.1. If a guide still leads with Sora, it is out of date.

Don't want to pick a model every time?

Because Quickads owns its tool and is not locked into any single model, it routes your brief to the right engine automatically: Veo for talking presenters, Kling for stylized cinematic shots, Seedance as the reliable default. You get the benefit of choosing the right model without doing the choosing.

Skip the guesswork, start free

The parts of a motion prompt

A good motion prompt is a short shot list, not a description. Five parts, in this order. Miss any one and you get one of the failures from the top of this guide.

  1. Cinematography

    Define the shot first: "medium shot", "close-up", "slow dolly in". This is your first creative decision, and it sets everything after it.

  2. Subject

    Who or what is the focal point. Name it plainly.

  3. Action

    What they are doing. Be specific about direction, speed, and intention. "Lifts the bottle toward camera and twists the cap" beats "uses the product".

  4. Context

    The environment and what sits in the background. This is where a scene stops feeling like a void.

  5. Style and ambiance

    Mood, lighting, lens, and finish. "Warm morning light, shallow depth of field, subtle 35mm grain" does far more work than "cinematic". Concrete beats vague every time.

That is the craft layer. For an ad, add three more decisions, covered next.

Pro tip. Ask your LLM to draft the motion prompt, and tell it which video model you are writing for, since the best phrasing differs by model. Many tools also have a prompt assistant built in. Search "prompt generator" inside the tool you use.

Direct for the outcome

Craft gets you a clip. Direction gets you an ad. Three more decisions turn a nice shot into something that performs.

  1. The hook, in the first second

    The opening frame decides whether anyone watches the rest. Open on motion or on the product mid-frame, not on a slow logo fade. Problem-first and pattern-interrupt openings tend to hold attention better than brand-first ones.

  2. The product truth

    Show the thing doing its job. A pump lifting to the applicator. A shoe flexing on a run. Demonstration beats description, on screen just as much as in copy.

  3. One clear action

    On-screen text and a single reason to act, timed so it stays readable. One ask, not three.

Relative stopping power by hook type

Directional benchmark from patterns we see, not a guarantee. Test against your own audience.

Problem-firsthigh
Pattern interrupthigh
Questionmedium
Bold claimmedium
Brand-firstlow

This is the steering. Volume and speed are the easy part now. You can generate a hundred variations. The teams that win point every one of them at the same outcome, then test which hook, which product moment, and which action actually move the number.

Build your prompt

Set the pieces below and this assembles a motion prompt you can copy straight into your model. It tunes the phrasing to the model you pick and reminds you what that model needs.

Direct your shot

Every field maps to one part of the structure above.

Your motion prompt

Weak prompt, directed prompt

Same idea, two prompts. The only difference is direction.

Describes a scene

"A woman using a skincare product in a bathroom. Cinematic."

Directs a shot

"Slow dolly-in on a matte-black serum bottle held in one hand, the cap lifting to reveal the dropper. Sunlit bathroom counter, soft morning light, shallow depth of field, subtle 35mm grain. Open on the bottle mid-frame in the first second. On-screen text reads 'Your 30-second glow' from second 1 to second 4, sharp and stable. 9:16 vertical."

Describes a scene

"A running shoe, dynamic and cool."

Directs a shot

"Low-angle tracking shot following a white running shoe as it flexes and pushes off wet pavement, water kicking up in slow motion. Overcast city street at dawn, cool blue light, shallow depth of field. Product fills the frame from the first frame. On-screen text 'Built to rebound' appears second 2 to second 4. 9:16 vertical."

Three formats worth studying. Watch the first second, the product moment, and the single action in each.

Talking-head UGC ad. A presenter, on-screen captions, and social proof. Notice the bold-offer hook and the single call to act.
Native in-car rant. Handheld and authentic, built for a problem-first hook that stops the scroll in the feed.
At-home talking-head testimonial. Soft natural light, straight to camera. The calm, personal format that builds trust before the ask.

The pre-generation checklist

Run this before you spend a single credit. Screenshot it and keep it next to your prompt window.

  • Directing, not describing. Camera, motion, and intent are all named.
  • Cinematography and style are concrete. A real lens and a real quality of light, not "cinematic".
  • Reference image locked. Image-to-video whenever the product has to stay consistent.
  • One small, motivated motion. No over-ambitious or self-contradicting movement.
  • Instructions do not conflict. Nothing in the prompt fights another line.
  • The ad layer is on. First-second hook, product truth, one clear action.
  • Aspect ratio matches the reference. Same format in and out.
  • The model is named. The prompt is written for the tool you are actually using.
Inside Quickads

Performance creative, as a service.

The raw models make footage. Turning that into on-brand creative that performs, across every format you run and at the volume modern media buying needs, is a different job. That is where Quickads sits. We are a performance creative as a service partner, a full content engine, not a single-format tool. Video is one output. So are motion banners, lifestyle banners, email graphics, landing and web pages, static and UGC ads, creator marketing, and the rest of what your calendar actually demands.

Motion banners

Animated display units, sized and looped for every placement.

Lifestyle banners

On-brand product-in-context stills, generated at catalog scale.

Email graphics

Header art and campaign visuals that match the rest of the funnel.

Web pages

Landing and product pages built to convert the traffic you buy.

Ads

Static, UGC, video, and clones of a winning creative, on repeat.

What makes that work. We own an award-winning proprietary tool, so we are never locked into a single model or stack. On top of it sits our own data and a team trained across every major third-party tool. That combination is built for flexibility, scale, and a skillset advantage you would struggle to hire for in one place.

Why it beats building in-house. The math is simple. The risk is close to none, the engagement flexes up and down with your calendar, and it runs at roughly half the cost of standing up the same capability internally. More on-brand creative, on more formats, for less.

Where these numbers come from

Real
30,000+
brands creating with Quickads
Real
30M+
ads in the Quickads ad library

A note on the figures here. The two numbers above are real. The performance guidance, which hooks tend to hold attention and what tends to convert, is directional: patterns we see, framed to help you form a hypothesis, not a promise. Model capabilities and availability are current as of July 2026 and change quickly. Validate everything against your own results before you rely on it.

The questions we keep getting

What is the best AI video model in 2026?

There is no single best model. Seedance 2.0 is the balanced starting point for most teams. Google Veo 3.1 leads on realism and synchronized dialogue. Kling 3.0 is best for stylized 4K and multi-shot work. Runway Gen-4.5 gives the tightest hands-on control.

Should I use text-to-video or image-to-video?

Use image-to-video for anything that has to look a specific way, especially product and ad work. Starting from a locked reference image keeps the look and the product consistent before the model animates a frame. Text-to-video is best kept for loose exploration.

How do I write a good AI video prompt?

Write it like a short shot list. Set the cinematography, name the subject, direct the action with specific speed and intention, set the context, and fix the style with concrete lighting and lens language. For an ad, add a first-second hook, show the product working, and give one clear action.

How long can AI video clips be?

Most models generate five to ten seconds in a single pass, and some now reach fifteen to twenty seconds. For longer videos you stitch clips together. Check the limit for the model you use, since it moves often.

How do you optimize AI video ads for conversions?

Optimize for the outcome, not the shot. Give the clip a hook in the first second, a clear moment of the product doing its job, and one reason to act. A pretty scene without those three will not convert.

Is it safe to use AI video ads commercially?

Licensing varies by model, so check each tool's terms. Adobe Firefly Video is trained on licensed content and is the conservative pick for regulated industries. Confirm rights before you run paid media.

Why does my AI video look generic?

Because you described a scene instead of directing a shot, and left the camera, lens, and light to the model. Name the cinematography and use concrete style language, then it starts to look like yours. See issue 01.

Why does the product change between shots in AI video?

Because you generated from text only, so the model reinvents the product every frame. Start from a locked reference image and state what must stay identical before you describe any motion. See issue 02.

Why does AI video motion look warped or melted?

Because the action you asked for was too big or contradicted itself, which is where models still fall apart. Keep the motion small and motivated, and remove any conflicting instructions. See issue 03.

You have the tools. Now direct them.

You can direct the shot now. Quickads produces it, on-brand and across every format, fast enough to keep testing, so your team can focus on what actually moves the number.