All articles
·Mike Serfort·9 min read

AI Video Prompts for Product Ads: Templates That Work (2026)

In short — A good AI video prompt for a product ad is one paragraph in a fixed order: camera, subject and action, environment, light, look. It ends with a short guardrail line (no text overlays, speaking or silent). Three rules matter more than clever wording. Never write your label or logo text into the prompt: let your product photo carry it. Describe motion, not the product, when you animate a photo. Repeat a continuity line from shot 2 onwards so three separate clips cut together as one ad. Below: the formula, the rules, and copy-paste templates for a 3-shot ad in both voiceover and talking versions.

Disclosure: Klaivo is our product. It writes these prompts automatically, and the rules below are the ones our production pipeline enforces on every ad across Veo 3.1, Seedance 2.0 and Kling 3.0. Model-specific guidance is linked to each vendor's own documentation.

Why most product-ad prompts fail

Prompts copied from AI video showcases are written for beautiful footage. Product ads fail in four specific, predictable ways:

  1. The label warps. The model redraws your packaging text and gets letters wrong. Shoppers notice immediately.
  2. The product drifts. Shot 1 shows your bottle; by shot 3 it's a slightly different bottle.
  3. The shots don't cut together. Three clips generated independently land in three different rooms with three different lighting setups.
  4. People talk when you didn't want them to (or mouth silently when you did). Audio behaviour wasn't specified.

Each rule below exists to prevent one of these.

The 5-part formula

Google's own guidance for Veo 3.1 structures a prompt as cinematography + subject + action + context + style & ambiance (Google Cloud). For product ads we use the same backbone, in this order, in English (video models follow English prompts most reliably, even when the voiceover is in another language):

# Slot What to write Example
1 Camera Shot size, movement, lens "Macro close-up, slow dolly-in, shallow depth of field"
2 Subject & action What happens, slowly and physically "a hand lifts the provided product bottle and turns it toward camera"
3 Environment One stable setting, with textures "on a light oak bathroom counter, white tiles behind"
4 Light Source and quality "soft morning window light from the left"
5 Look The ad style "clean premium e-commerce look, natural colors"
+ Guardrails Text and audio behaviour "clean frame, no text overlays, no watermarks, silent motion"

Put together:

Macro close-up, slow dolly-in, shallow depth of field: a hand lifts the provided product bottle and turns it toward camera, on a light oak bathroom counter with white tiles behind, soft morning window light from the left, clean premium e-commerce look with natural colors. Clean frame, no text overlays, no watermarks, silent motion, subject has closed lips.

Rule 1: never write your label text into the prompt

This is the single most valuable rule for product ads, and the least obvious. If you write "a bottle labeled GLOWLAB Vitamin C Serum", the model tries to render those letters, and AI video models are unreliable typographers.

Instead:

  • Give the model your real product photo (image-to-video) and let the image carry the label.
  • Refer to the product generically: "the provided product bottle", "the jar", "the item on the table".
  • Ask for motion that keeps the label readable: slow rotations, the label facing camera, no fast spins.

We apply this on every generation, and it's also why image-to-video beats text-to-video for e-commerce. See image-to-video for product ads for the full reasoning.

Rule 2: when animating a photo, describe motion, not the product

In image-to-video, the model already sees your product. Re-describing its shape, color and material competes with the image and invites the model to "correct" it. Spend your words on what the image can't say:

  • Camera movement: dolly-in, orbit, tilt up, handheld drift.
  • Action: a hand enters, the cap twists off, liquid pours.
  • Light change: a reflection slides across the glass, sunlight warms the scene.

A useful test: if a sentence in your prompt would still be true of the still photo, delete it.

Rule 3: add a continuity line from shot 2

Each shot of a 3-shot ad is a separate generation with no memory of the others. Without instruction, you get three rooms. From the second shot on, open the prompt with an explicit continuity line:

Visual and light continuity with the previous scene, in the exact same bathroom with light oak counter and soft morning window light…

Two habits make this reliable:

  • Choose one setting for the whole ad and repeat its key nouns word for word (same counter, same tiles, same light direction).
  • Keep the style slot identical across all three shots. Changing "premium e-commerce look" to "cinematic look" mid-ad is enough to break the cut.

Rule 4: decide speaking or silent, explicitly

Modern models generate audio, and Veo 3.1 always does (Gemini API docs). Leave audio unspecified and you'll get improvised mumbling or mouths moving under your voiceover. Pick one mode per ad:

Voiceover ads (product is the hero). End every prompt with:

silent motion, subject has closed lips, calm expression, no speaking.

Talking ads (person is the hero). Put the exact line in quotation marks, in the target language (quotation marks for dialogue are also Google's documented convention):

…a woman in her late 20s speaks naturally to camera with realistic lip-sync. She says out loud in English, clearly and naturally: "I stopped buying three serums the day I tried this one."

Keep spoken lines to 8–12 words per 5-second shot, which is what fits. The full reasoning is in the video ad structure that sells.

Rule 5: use style modifiers that describe light and camera, not adjectives

"Stunning" and "high quality" do nothing. Styles work when they name camera behaviour and lighting. The modifiers we use:

Style Modifiers that actually change the output
UGC hand-held smartphone camera, subtle natural movement, soft daylight by a window, everyday apartment, natural tones
E-commerce close-up detail shots, studio softbox key and fill light, neutral clean backdrop, focus on texture and use
Luxury slow motion, rim light and backlight, delicate specular reflections, minimalist set, deep shadows
Cinematic anamorphic lens look, high-contrast side lighting, shallow depth of field, volumetric light rays

Copy-paste templates: a 3-shot product ad

Replace the bracketed parts. Keep everything else, especially the continuity line and the guardrail tail.

Voiceover version (product-first)

Shot 1 — hook

Macro close-up, fast push-in with a subtle speed ramp: [a hand grabs the provided product] and brings it toward camera, [on a clean kitchen counter], [warm morning light], [clean premium e-commerce look]. Clean frame, no text overlays, no watermarks, silent motion, no speaking.

Shot 2 — benefit

Visual and light continuity with the previous scene, in the exact same [kitchen counter] setting. Medium close-up, slow orbit: [the product in use — e.g. powder dissolving in a glass of water], [warm morning light], [clean premium e-commerce look]. Clean frame, no text overlays, no watermarks, silent motion, no speaking.

Shot 3 — call to action

Visual and light continuity with the previous scene, in the exact same [kitchen counter] setting. Slow dolly-in to a final packshot: the provided product standing upright, label facing camera, [warm morning light], soft glossy reflections, [clean premium e-commerce look]. Clean frame, no text overlays, no watermarks, silent motion.

Voiceover script alongside (8–12 words each): "Your afternoon crash isn't about sleep." · "One scoop, no sugar, steady energy for hours." · "Try it for 30 days — link below."

Talking version (person-first)

Shot 1 — hook

Vertical selfie-style shot, hand-held smartphone camera with subtle natural movement: [a woman in her late 20s] in [a bright bathroom] looks straight into camera with a surprised expression, [soft daylight from a window], authentic UGC look. She says out loud in [English], clearly and naturally: "[Hook line, 8–12 words]"

Shots 2 and 3 use the same person, room and light, open with the continuity line, and carry one quoted line each: the benefit, then the call to action.

Model-specific notes

  • Veo 3.1. Clips are 4, 6 or 8 seconds, and 1080p or 4K only at 8 seconds. Google documents timestamp prompting ([00:00-00:02] … [00:02-00:04] …) to fit several beats in one clip, useful for a fast hook (Google Cloud). For exclusions, Google advises describing them concretely as part of the scene ("a desolate landscape with no buildings or roads") rather than as a bare negative ("no man-made structures").
  • Kling 3.0. Up to 15 seconds per clip with multi-shot generation and multi-character dialogue (Kling). It's a good fit when one long talking shot replaces three short ones.
  • Seedance 2.0. Accepts several reference images, videos and audio clips in a single request (ByteDance). With one product photo, the rules above apply unchanged.

Which model to pick for which shot, with costs, is covered in Seedance 2.0 vs Kling 3.0 vs Veo 3.1 for video ads.

Pre-flight checklist

Before you generate, check that the prompt:

  • follows camera → subject & action → environment → light → look;
  • contains no brand name or label text;
  • calls the product "the provided product…" when a photo is attached;
  • names one setting, repeated word for word in every shot;
  • opens shots 2 and 3 with a continuity line;
  • states the audio mode (silent with closed lips, or a quoted line);
  • keeps each spoken line to 8–12 words;
  • uses light and camera words instead of empty adjectives.

FAQ

Should AI video prompts be written in English?

Yes, for the visual prompt: models follow English camera and lighting vocabulary most reliably. The spoken line or voiceover can be in any language the model or voice supports. State that language explicitly inside the prompt.

Why does my product label look wrong in AI video?

Usually because the label text is written in the prompt, or because the product was generated from text instead of a photo. Remove the text, attach your real product image and ask for slow movement that keeps the label facing camera.

How long should a video prompt be?

Long enough to fill the five slots, typically 40–80 words per shot. Beyond that, extra adjectives rarely change the output and can dilute the instructions that matter.

Can I use the same prompt on Veo, Kling and Seedance?

The structure transfers well. Adjust length to each model's clip duration (Veo's 8 seconds versus 5 on Kling and Seedance) and remember that Veo always generates audio, so voiceover ads need the silent guardrail.

Do I have to write prompts myself?

No. Tools like Klaivo write the script and the shot prompts from a product page or photo, applying these rules automatically. The rules above are still useful when you want to edit a shot by hand.

Want to try it on your product?

Create your first ad