Last week a TikTok Shop seller sent me a product photo of a moisturizer at 11pm and asked if I could "just make the ad." I didn't. I screen-shared Vidu Q3 with her, we built three versions together in about 20 minutes, and she picked one to post by morning.
That's the real job this tool does — not "AI makes your ad for you," but shortening the path from a clean product photo to a postable short-form clip you can test. This piece walks through the workflow I actually use, the reference setup that works, the prompt structure that doesn't waste credits, and where it reliably breaks. If you run dropshipping, affiliate, or TikTok Shop creative, this is the one-hour version of everything I've learned from testing.
I'm Maya. Follow me now.
What this Vidu Q3 workflow is best for
Before we touch the tool, be clear about the job.
Vidu Q3 Reference-to-Video does one thing the rest of the AI video stack still struggles with: keeping your product visually intact while the scene around it moves. You feed it 1–4 reference images, describe the scene briefly, and get a 10–16 second clip with native audio. The launch materials from ShengShu position it for advertising and e-commerce directly, which is unusually honest for a model release post. Third-party benchmarks back the quality claim — Vidu Q3 currently sits at No. 2 globally on the Artificial Analysis video leaderboard behind Sora 2.

Practically:
- Works well for: product promo clips, catalog variations, scene A/B tests, multi-language ad localization, "how it looks in the hand" lifestyle shots.
- Works badly for: consistent brand mascots across a multi-week campaign, ads with fine legible text on packaging, cinema-grade hero pieces needing specific directorial control.
Everything below is tuned for the first list.
What to prepare before you generate
The single biggest reason first clips look bad: people skip prep and burn credits figuring it out. Ten minutes of prep saves 20 minutes of regen.
Product images and references
You get up to 4 reference slots. Don't use all 4 unless you have a reason. Two is the sweet spot.
What makes a strong product reference:
- Clean background. White, neutral gray, or minimally busy. Complex backgrounds bleed into generation.
- Sharp, even lighting. Soft daylight beats hard studio light for this model.
- The angle you want in the final clip. Vidu animates from roughly this framing. If your reference is a top-down flat lay but you want a hand-held shot, the model has to reinterpret — and reinterpretation is where things drift.
- High resolution. At least 1024px on the short edge. Sharp phone photos work.
What to avoid: stock photos with visible watermarks, compressed listing screenshots (compression shows up as shimmer), multiple products in one frame.

Visual style cues
Your second reference (if you use one) should set the vibe, not the product. A mood shot. A competitor ad frame. A Pinterest save. Think "this is the world I want my product to live in."
Common mistake: picking a style reference visually incompatible with your product — "luxury spa" style for a $9 dropshipping gadget. The model won't fight you, but the output will feel off.
Hook and scene intent
Before you open the prompt box, write one sentence for each:
- The hook. What's the first second doing? ("Hand reaches into frame, picks up product.")
- The middle. What's the product doing? ("Camera slow-dollies in as product is applied to skin.")
- The close. How does the clip end? ("Cut to wide shot with logo visible on bottle.")
These three sentences become your prompt skeleton. Writing them before generation is non-negotiable if you want consistency across hook variations. I've seen people skip this and generate 10 clips that all feel vaguely the same because they kept prompting from vibes.
Step-by-step: build a product clip with Vidu Q3
Open the Vidu app, select Reference-to-Video, and follow this.

Set the subject and scene references
Upload your product photo first — this is your subject. Upload one style/mood reference second. If you need a talent reference (face for a UGC-style read), that's third. Don't upload all four on your first pass.
Label them in your head by role: Product = visual fidelity. Mood = world/lighting. Talent = face/body. Style = optional aesthetic.
Add motion, camera, and pacing instructions
Here's where people write too much or too little. The prompt should describe what's visible, not what's emotional.
Good prompt:
"Hand reaches in from bottom frame, picks up the product, slowly tilts it toward camera. Camera holds steady, soft natural light from window, morning kitchen setting. Ends with product held at eye level, subtle smile."
Bad prompt:
"Luxurious and empowering moisturizer ad that feels aspirational and makes viewers want to buy the product for their self-care routine."
The first gets you a usable clip. The second gets you something that looks like every other AI ad on the platform.
On camera control: describe physical movement in plain language. "Slow push-in on product," "static medium shot," "follow hand as it moves." ShengShu's technical overview lists six kinds of cinematic effects the model handles, but the pattern is simple: simple camera moves land consistently, complex ones are aspirational.

Layer in audio intent and dialogue cues
This is where Vidu Q3 pulls ahead — you specify audio in the same prompt, generated in one pass.
- Voiceover: write the exact line in quotes. "Say: 'This is the one I wish I'd found sooner.'" Multilingual works — write in Spanish or Japanese and get appropriate lip-sync.
- Sound effect: describe it. "Soft click as cap opens, gentle ambient kitchen sounds."
- Background music: tone, not genre. "Warm, upbeat, low-key." Don't name specific songs.
Honest caveat: audio nails it about 70% of the time. The other 30% you regen or swap in post. Brand names with unusual spelling get mispronounced.
Generate multiple hook variations

This is the workflow that actually saves time, and it's the step beginners skip.
Lock your product photo. Lock your style reference. Keep everything the same except the first 1–2 seconds of the scene description. That's your hook variation.
Example — same product, three hooks:
- "Hand picks up product from counter, tilts toward camera."
- "Close-up on product, pulls back to reveal kitchen scene."
- "Product sits on counter, hand enters frame, picks it up with slight excitement."
Three generations, roughly 10 minutes, three hook options for testing. TikTok's creative guidance reinforces what every operator knows — the first 1–2 seconds do most of the CTR work.
Run all three at 720p first (save credits). Pick the winner. Regenerate that one at 1080p with audio for final output.
Common mistakes and how to fix them
Six patterns I see on repeat.
- Using all 4 reference slots on attempt one. More references = more conflicts for the model to resolve = weaker output. Start with 2. Only add references if the first pass is missing something specific.
- Writing emotional prompts instead of visual ones. "Feels aspirational" doesn't generate. "Camera dollies in, natural window light, soft colors" generates. Describe what the camera sees, not what the viewer should feel.
- Skipping the 720p draft. People generate at 1080p with audio on attempt one, then regenerate three more times when the composition's off. That's 4x the credit cost for a clip you could've locked at 720p first.
- Stacking incompatible references. Luxury spa mood + budget dropshipping product = off-brand weirdness. Make sure your references belong together.
- Over-directing camera movement. "Whip pan left, crash zoom, reverse angle cut" is ambitious. The model ignores most of it and picks something close. One camera move per clip, maximum two. Keep it simple.
- Ignoring the product's fine details. If your product has small legible text or fine pattern, add motion cautiously — details shimmer under heavy camera work. Static hero shots handle detail better than dynamic product-in-motion shots.
Limits and trade-offs to know before publishing
Blunt list. Don't hit publish without reading this.
- Consistency across separate generations is weak. The same face or product across three separate runs will drift slightly. For one clip, the clip is internally consistent. For a three-ad campaign featuring the same digital talent — pick a different tool.
- 16 seconds is a hard cap. If your ad needs 30 seconds, you're stitching two generations, and the seam shows. Vidu is designed for the short-form slot specifically.
- Product text degrades under motion. If your packaging has fine print that has to stay legible, do a test generation first. If it morphs, use a static product angle and move the camera instead of the product.
- Audio regen costs stack up. The voiceover will be wrong sometimes. Plan to either accept it, regen once (costs credits), or swap in a real voice in post.
- Too-clean output reads AI. This is a platform problem, not a tool problem. Audiences on TikTok and Reels are pattern-matching "made with AI." If your clip looks too polished, it may underperform authentic iPhone footage even if it's technically better. Counter this by keeping lighting natural, avoiding overly cinematic camera moves, and sometimes adding subtle grain in post.
None of these are dealbreakers. All of them are real. Knowing them up front saves you from publishing something that underperforms for reasons you didn't see coming.
FAQ
How many references should you use?
Two in most cases. Product + mood. Three if you need a specific talent face. Four only if you have a clear role for each reference and you've already tested with fewer. More references doesn't mean better output — it means more constraints the model has to juggle, and juggling is where drift happens.
Can Vidu Q3 make ecommerce product videos?
Yes, and this is one of its strongest use cases. Ecommerce product video is exactly the brief: product stays visually intact, scene animates around it, camera moves modestly, 10–16 second output lines up with Amazon video placements, TikTok Shop creatives, and Shopify product page embeds. The workflow above is built around this job specifically.
Is this workflow good for TikTok Shop creatives?
For paid TikTok Shop ads, yes. For organic TikTok Shop content, with reservations. Paid ads benefit from polished production; the "produced" quality of Vidu output works there. Organic TikTok Shop content does better with raw, creator-style footage — audiences trust it more. A reasonable split: use Vidu for your paid testing fleet, use a real creator (or your phone) for organic.
Do you still need manual editing after generation?
In most real workflows, yes. Typical cleanup: trim the first half-second if it starts slow, retime subtitles, sometimes overlay your own hook text at the start, occasionally swap the AI voiceover for a better take. CapCut handles all of it in under 5 minutes per clip. "Zero edit" is marketing copy. "80% less editing than shooting from scratch" is honest.
Related Articles

Wan 2.1 Image-to-Video Prompting Guide
Learn how Wan 2.1 image-to-video workflows can support short-form clips, prompt control, and creator-friendly motion tests.

Maya
Jul 8, 2026

Viyou Alternatives for AI Video Inspiration
Explore Viyou alternatives for AI dance videos, image-to-video clips, and short-form creative inspiration workflows.

Maya
Jul 8, 2026

Vidnoz Image-to-Video Review for Social Clips
Is Vidnoz image-to-video useful for social clips? This review looks at workflow fit, limits, pricing, and short-form creator use cases.

Maya
Jul 8, 2026

Vheer AI Image-to-Video Review for Social Clips
Is Vheer AI image-to-video useful for social clips? This review looks at workflow fit, output limits, and creator use cases.

Maya
Jul 8, 2026

