Which shots hold up on camera, which still fail, and the one check that catches each category's worst failure. Every page on this subject is written by someone selling AI creative or selling the alternative. We sell one of them, so these limits are drawn against our own interest.
Stated up front, because the whole value of this document depends on it.
The same model that renders a flawless mid-shot of a face will fuse the fingers when a hand comes near it. Quality is not a property of the tool, it is a property of the shot you asked for, which means the decision belongs at shot level in the storyboard, not at tool level in a procurement meeting.
The recurring failure is an effect with no physical cause: fingers hovering rather than gripping, feet not meeting the floor, a cheese pull that stretches forever because nothing tells it to break. Viewers detect this before they can name it, which is why "it looks fine but feels wrong" is the most common review note in the category.
Ask for a push-in and a subject doing something in a five-second clip and most models silently pick one. The move becomes a drift, or the action stalls. This is the most useful single constraint we have found, because it is predictable and it is cheap to design around.
Jewelry is prong count. FMCG is front-of-pack under four lighting conditions. Food is char and the cheese break. Beauty is skin around the mouth and whether the bullet holds its shape once used. Fashion is scale against a body. None of these are generic quality checks, and none are in any tool's default QA.
Our own warp detector measures up to 23% on clips that are completely fine. Hand and face distortion detectors behave similarly. They are useful as a warning that directs a human to a timestamp; they are not useful as a gate. Anyone selling automated AI-artifact detection as a pass/fail decision is overselling it, and that includes anyone selling ours that way.
Worth stating first, because the failure list is longer and gives a misleading impression. These are shots we now run without a second thought.
| Shot | Why it holds | Watch for |
|---|---|---|
| Product silhouette across a cut | Shape is the strongest signal a model carries between frames | Drift on reflective or faceted objects |
| Brand colour, single lighting setup | One condition means one value to hold | Colour walks when the setup changes mid-cut |
| Face, single subject, mid-shot | The most-trained subject in every model | Identity drift across separately generated clips |
| Skin at conversational distance | Texture reads correctly until the camera gets close | Plastic skin in macro; pores are the tell |
| Front-of-pack, static angle | A flat surface held still is close to an image task | Falls apart the moment the pack is tilted |
| Talking head under 15 seconds | Short duration limits how far anything can drift | Lip sync and dead eyes |
| Fabric drape, mid and wide | Cloth behaviour is well modelled at distance | Seams and stitching in close-up |
| Environmental b-roll | No product fidelity requirement, no hands | Signage and text in the background |
Everything that holds is either a single subject, a single lighting condition, a short duration, or a shape rather than a mechanism. Everything in the next section violates at least one of those. That is the whole heuristic, and it predicts new failures better than any list of known ones.
Seven recognisable classes, taken from the failure taxonomy in our own clip QC. They recur across models and across categories.
| Failure class | What it looks like | Cost when missed |
|---|---|---|
| Action without physical cause | Fingers hovering rather than gripping, feet not meeting the floor, contact that never happens | Viewer feels it before naming it. The ad underperforms and nobody can say why. |
| Localised render artifacts | Fused fingers, a colour blob in a hand, melted geometry at a join | Brand review rejection, usually at the last approval stage |
| Continuity breaks | Object count changes between shots, a label rewrites itself, a garment changes | In jewelry and FMCG this makes it a different product, not a different shot |
| Prop leaves and never returns | The acting object drifts out of frame mid-clip and the shot stops reading | The beat no longer shows what it was written to show |
| Prompted direction ignored | A push-in that is really a drift, a slow move that never starts | The board and the cut no longer match. Usually found in edit. |
| Prompted detail absent | The steam, the condensation, the dust that was the entire texture of the shot | Technically passes every automated check |
| Composition against placement | Subject sitting in the 9:16 caption band, half the frame dead | Fine in the preview, unusable in the placement |
Counter-intuitive and consistent. A face is the most-trained subject in any model; a full-body run across a room is a coordinated mechanical sequence with a dozen joints that all have to agree. If a board calls for someone sprinting, budget re-runs for it and do not budget them for the close-up.
Six months of production, reduced to the single check that kills more work than everything else in that category combined.
| Category | The check | Why this one | Still shoot |
|---|---|---|---|
| Jewelry | Count the prongs in every shot | A prong that migrates between frames means it is not the same ring, and in jewelry the ring is the entire product. Then the shank, where band meets setting, and that join softens. Then the reflection, which has to agree with the object above it. | Diamond fire. True dispersion renders as generic white sparkle. For a solitaire hero at full screen, shoot it. |
| Beauty | Does the skin around the mouth still read as skin | Then: does the case hold the same colour in macro and in wide, does the shade hold under two lighting setups, does the bullet keep its shape once used. | Any shot with fingers near the face. It is the most re-run shot we have. |
| FMCG and packaged goods | Would it pass as the pack you would pick up in a shop | Not whether it looks good. A pack has a fixed logo lockup, colour, front photograph, weight and tier. Slightly wrong is not an ad, it is a counterfeit that brand review kills and a shopper does not recognise. | Pours, scatters and anything granular. Grains moving as a clump is the giveaway. |
| Food | Char, basil, and whether the pull breaks | Real ovens leave uneven blistering and near-burnt spots; generation smooths to an even golden ring that reads as plastic. Leaves wilt where it is hot and stay flat where it is not. Cheese stretches, thins and breaks. A pull that goes on forever is animation. | Full-body movement. Faces are easier than legs. |
| Fashion and accessories | Scale against the body | Bag and accessory sizing is the hardest solved problem in the category and the one generic tools get most visibly wrong. A bag that reads one size in the wide and another in the close-up is the commonest fashion tell. | Seams, stitching and hardware in macro. |
| Supplements and regulated | Is every word on the label legible and correct | A label carrying printed claims, on screen and readable, is both a fidelity problem and a compliance one. Wrong text on a regulated pack is a different category of mistake. | Anything where the claim itself is the creative. |
Every one is checkable by anyone in ninety seconds, against a clip they did not make. That is deliberate. A limitation you can verify is worth more than a capability you have to take on faith, and it is the only reason to believe the rest of this document.
The most useful constraint we have found, because unlike most limits it is predictable, and designing around it costs nothing.
Three stages. The first two are arithmetic and can be automated. The third needs a person looking at frames, and it is the one that catches the faults that cost a reshoot.
| Stage | What it catches | Automatable |
|---|---|---|
| A · Technical | Resolution downgrades, black or frozen frames, chroma blowout, colour jumps, silent or under-level audio | Fully |
| B · Brand | Palette conformance, logo presence, safe zones, banned and required wording: the part of a guideline that lives in pixels | Fully |
| C · Frames | Whether the shot shows what it was written to show. Contact, continuity, camera move, absent detail, composition against placement | Not at all |
A model quietly returning 720p when you paid for 1080p is caught at the clip stage rather than discovered three days later. Worth automating for this alone.
Stage A measures pixels. It has no idea what the shot was for, so it will pass a clip that does not depict the thing it was written to depict. A frame pass where nobody looked at a frame is not a frame pass.
Twelve frames in a grid with timestamps burned in. The question that finds the expensive faults: is anything present early and gone later, or the reverse.
The AI-artifact detectors (warp, hand instability, face distortion) are uncalibrated. Ours measures warp of up to 23% on clips that are known to be fine. They warn; they do not block, and we do not let them trigger a regeneration on their own. Any vendor presenting automated artifact detection as a clean pass/fail is describing something that does not currently exist.
A finding is half the job. The change is usually one clause in the prompt, not a new brief.
| What went wrong | What to change |
|---|---|
| Artifacts in hands or faces | Reframe away from the problem area, or switch to image-to-video from a still that is already clean. Reduce how much is moving at once. |
| Action has no physical cause | Name the contact and its consequence, not the outcome. Then drive from an approved first frame rather than from text. |
| Prop leaves frame and never returns | Lock the opening geometry with a generated still showing correct contact, pass it as the first frame, and add an explicit persistence clause. Consider a shorter duration. |
| Camera move ignored or reversed | Drop the move or extend the duration. If the move matters, make it the only motion in the prompt. |
| Subject in the caption band | A reframing instruction is usually enough. Changing aspect ratio is the heavier option. |
| Product colour or label drifting | Fewer lighting conditions in one cut. Every change of setup is another chance for the SKU to walk. |
Rewriting six clauses at once means the next clip tells you nothing about which one mattered. This is the most commonly ignored rule and the most expensive.
When a shot keeps failing, generating a clean still first and animating from it solves a surprising share of hand, face and product-fidelity problems in one move.
Uncalibrated detectors are not grounds to spend money. Look at the frame the warning points to, then decide.
The useful question is not whether to use AI. It is which four shots on the board still justify a camera, and the answer is consistent enough to plan against.
It does not remove the shoot. It changes what the day is for. A day spent capturing four hero shots that generation cannot do, which then feed a hundred generated variants, is a different proposition to a day spent capturing a hundred shots at a hundred-shot cost. That is the honest version of the argument, and it is more useful than either the replacement claim or the dismissal.
Two moves per phase. Nothing here needs new software.
Everything here describes September 2026. Hands were meaningfully worse a year ago and will be better a year from now. The durable part is not the list. It is the heuristic underneath it: single subject, single lighting condition, short duration, shape rather than mechanism. Anything violating one of those is where to point the review.
Research winning ads. Generate high-converting creatives. Ship them at volume. All under one roof. The limits in this report come from six months of solving them, per category, on our own production.
Ads in Library
Languages
Brands Trust QuickAds
Search and analyse competitor ads across Meta, Google and TikTok. Filter by industry, format, hook type, engagement, brand and market. The 32M+ library is updated continuously.
Generate high-performing ad creatives using AI: static, video, carousel. 1200+ creatives a month per brand, five to seven days from brief to first full batch. AI handles form; humans own substance.