Sarah Johnson

Tutorials

How to Write AI Video Prompts That Actually Work (With Examples)

How to write AI video prompts that actually work

The best AI video prompts follow a five-part structure: subject, action, setting, camera, and lighting — in that order, written as one dense paragraph. Master that structure and any generator, from Veo 3 to Kling, produces usable footage on the first or second try instead of the tenth. This guide gives you the exact framework plus copy-paste examples for the most common ad shots.

Why do most AI video prompts fail?

Most prompts fail because they describe a vibe instead of a shot. "A cool video of my skincare product" gives the model nothing to work with — it has to guess the subject's position, the environment, the camera behavior, and the mood, and it will guess differently every generation. The result is the frustrating loop most beginners know: ten generations, ten wildly different clips, none usable.

Professional prompts remove guesswork. They read like instructions to a cinematographer: what is in frame, what it is doing, where it is, how the camera moves, and how the scene is lit. The model still adds creativity, but inside boundaries you set.

What is the five-part prompt structure?

Every reliable video prompt answers five questions in sequence:

1. Subject — who or what is on screen, described physically. Not "a woman" but "a woman in her late 20s with shoulder-length brown hair, wearing a cream knit sweater." Specificity locks identity across generations.

2. Action — one clear motion. AI video handles a single continuous action far better than a sequence. "She lifts the serum bottle and presses the dropper" beats "she does her skincare routine."

3. Setting — the environment with two or three concrete details. "A bright bathroom with white subway tile and a small potted plant on the counter" renders consistently; "a nice bathroom" does not.

4. Camera — shot type and movement. "Handheld selfie-style, slight natural shake" or "slow push-in from a tripod at chest height." If you skip this, the model picks for you, and it usually picks wrong.

5. Lighting — the mood anchor. "Soft window light from the left" or "warm golden-hour glow" or "bright even ring light." Lighting language does more for realism than any other single element.

What does a complete prompt look like?

Here is a full UGC-style example you can adapt: "A woman in her late 20s with shoulder-length brown hair, wearing a cream knit sweater, holds a small amber serum bottle up to the camera and taps the dropper twice. She is in a bright bathroom with white subway tile and a small potted plant on the counter. Handheld selfie-style camera at arm's length with slight natural shake. Soft morning window light from the left. She speaks casually to the camera with a relaxed smile."

Notice it is one paragraph, present tense, and dense with physical detail — no filler words like "beautiful" or "high quality," which add nothing the model can act on. This single-paragraph dense format is what current models are trained to follow best.

How do prompts change between models?

The structure stays the same, but each model rewards different emphasis. Veo 3 responds strongly to audio cues — add "she says: 'I was skeptical at first'" and it will generate synced speech. Kling excels at physical product motion, so lean into material detail: liquid pouring, fabric moving, steam rising. If you have not picked a model yet, our comparison of Veo 3 vs Kling vs Sora for ads breaks down which fits which job.

What are the most common prompt mistakes?

Four mistakes cause most failed generations. First, stacking multiple actions — split them into separate clips and edit together instead. Second, negative phrasing: models handle "empty counter" better than "no clutter on the counter." Third, vague adjectives like "amazing" or "professional" that carry zero visual information. Fourth, forgetting aspect ratio — always state vertical 9:16 for TikTok and Reels content, or you will get widescreen footage you cannot use.

How do you iterate when a generation is close but not right?

Change one variable at a time. If the subject looks right but the scene feels flat, touch only the lighting line. If the motion is wrong, rewrite only the action. Rewriting the whole prompt resets everything the model got right. Save your winning prompts in a document — a personal prompt library is the fastest compounding asset in AI content, because a proven prompt plus a small tweak is a new ad in minutes.

FAQ

How long should an AI video prompt be?
300 to 600 characters of dense, specific description is the sweet spot for most models — long enough to control the shot, short enough to stay coherent.

Should prompts be written in present tense?
Yes. Present tense reads as a scene description, which is how video models are trained. "She lifts the bottle" outperforms "she will lift the bottle."

Can I reuse one prompt across different products?
Yes — build a template with your structure and swap the subject and product details. That is how teams batch dozens of ad variations a day.

Do I need different prompts for images vs video?
The subject, setting, and lighting sections carry over directly; video adds the action and camera-movement lines. Write the image version first, then extend it.

Start Creating With SMPL Social

Everything in this guide can be done in your browser with SMPL Social — turn a product photo or a simple idea into AI video ads, UGC content, and scroll-stopping visuals in seconds. No filming, no editing, no credit card. Sign up free and get 50 credits to make your first piece of content today.

Sarah Johnson

Join the newsletter

Be the first to read our articles.

Follow Social Media

Follow us and don’t miss any chance!

Join the newsletter

Be the first to read our articles.

Follow Social Media

Follow us and don’t miss any chance!