Most "how to use GPT Image 2" posts I've read this week are either prompt dumps or lab demos. Neither helps if you're sitting on 30 product photos that need to ship as ad creatives by Friday. This one is for that situation.
I'm Maya. I'm going to walk through how I've been using GPT Image 2 on actual product visuals — what to prepare before you open ChatGPT or the API, how to structure the prompt so you don't waste 20 generations landing on something usable, how to produce multiple creative angles without starting over each time, and where to stop because the model hits its limits. No prompt templates you'll paste and abandon. The actual working pattern. This is for ecommerce ops, TikTok Shop sellers, affiliate and dropshipping operators, and small social ad teams making 5-30 testable variations — not brand hero images.
Why GPT Image 2 is interesting for product ad visuals

Before this model, the product-ad workflow had one recurring failure point: the text on the image. You'd generate a decent lifestyle shot, then rebuild the caption, price tag, hook copy, and CTA in Canva.
GPT Image 2 collapses that step. The model can draft the lifestyle scene and embed the tagline, price, or hook directly into the image in one pass. Is it perfect? No. But it's close enough that the finishing pass goes from "build from scratch in Canva" to "touch up a near-final image." That's the whole reason this is worth a tutorial.
The OpenAI prompting cookbook lays out the patterns that actually work in production: ads, infographics, UI mockups, photorealism. This tutorial follows those patterns but narrows them to the product-ad use case, because most operators don't need 12 patterns — they need the one that ships.
What to prepare before generating
The single biggest reason people waste time on this model is opening it cold with a vague prompt. Before you generate anything, have these three things ready in a single doc or folder.
Product photos
You want clean, well-lit product shots on plain backgrounds. Not the Amazon listing image with the watermark. Not the Instagram shot with a person half in frame. The cleaner the reference, the more control you have.
What works: a front-facing product shot, ideally one with subtle ambient lighting rather than a flat white cutout. If you have multiple angles, pull 2-3 — the model can use multiple references in a single call to composite better. If your only asset is a low-res catalog image, that's your ceiling. The model preserves input fidelity — a bad input stays bad.
One honest limit: per the OpenAI image generation guide, GPT Image 2 doesn't support transparent backgrounds. If you need a PNG with alpha for compositing later, you're going back to a different model or doing the cutout manually. Plan around it.
Offer message and hook
Write the exact text you want on the image before you prompt. Not "a tagline about saving money" — the literal words. Short hooks work best. Things like "Sold out 3x this year" or "Now $24, was $38" or "Before bed, every night." Not paragraphs.
This matters more than it sounds. The model's text rendering is strong now, but it's strongest when you give it 3-8 words and a clear hierarchy — headline vs. subhead vs. CTA. If you hand it a 30-word paragraph, expect kerning issues and broken line breaks. Keep it punchy, the way you'd write a hook for a Reel.
Brand style references
One or two images that represent the visual tone you want. Not a Pinterest board. One example of a product ad whose lighting, color palette, and mood you want to replicate. You'll feed this in as a style reference.
This is what separates outputs that look like stock AI content from outputs that look like they belong to your brand. Without a reference, you get the default "AI product scene" look — clean but generic. With one, you're telling the model "match this."
Step-by-step: create product ad visuals with GPT Image 2

Build the prompt structure
The structure that works, based on OpenAI's own cookbook and on actual testing: Scene → Subject → Key details → Constraints → Intended use.
Concretely, for a product ad, that's:
- Scene: where the product lives. "Morning kitchen counter, soft window light, light wood surface."
- Subject: what the product is. "8oz amber glass skincare bottle, labeled Restore Serum."
- Key details: textures, materials, supporting props, text elements. "Condensation on the bottle, small eucalyptus sprig to the left, headline text reading 'Sold out 3x this year' in clean serif."
- Constraints: what the model should not do. "No extra product, no human hands, no watermark, no background clutter."
- Intended use: "Vertical 9:16 social ad for Instagram Reels."
That last line matters more than most people realize. Telling the model the asset is for a social ad changes how it handles composition, text hierarchy, and negative space. Without it, you get a "nice image." With it, you get something formatted for the surface where it'll run.
One pattern I use constantly for edits: Change / Preserve / Constraints. If you already have a generated image and want a variation: "Change: background to a wooden bathroom counter. Preserve: product label, text overlay, lighting direction. Constraints: no new props, no style drift." This pattern saves an insane amount of time versus starting from scratch.
Generate multiple creative angles
Here's where most tutorials go wrong. They stop at one prompt. Stop doing that.
Run the same product through 3-5 different scene contexts in parallel. Keep the subject and text identical, change the scene. Same serum bottle but: on a morning counter, in a gym bag, next to a laptop, in a hand on a beach. Each one tests a different audience angle — wellness, fitness, corporate, travel.
The gpt-image-2 API on Replicate supports up to 10 images per call, with control over aspect ratio (1:1, 3:2, 2:3) and quality (low/medium/high/auto). For drafting, use quality: low — faster, cheaper, good enough to judge direction. Bump to high only for the variants you want to ship.

The mental shift: you're not making one ad. You're making a test set. Five scenes × three hooks = fifteen assets. Maybe three work. That's the game.
Evaluate text and layout quality
Don't judge outputs emotionally. Have a checklist. Open each generation and check:
- Is the text readable at thumbnail size? If it's illegible in the Instagram feed preview, it's dead. Zoom out before you judge it zoomed in.
- Is the hierarchy right? Headline should dominate. Subhead supports. CTA is clear but not competing. If your eye goes to the wrong thing first, the composition is wrong.
- Does the product look correct? Color accuracy, label accuracy, proportions. The model occasionally drifts on product details even with reference images. Check every time.
- Are there clipping issues at the edges? CTAs that get cropped, logos that bleed off — especially bad at 9:16. Sanity-check edges before approving.
- Does the text match what you specified? Generated text still occasionally hallucinates characters or word breaks, especially in longer phrases. Read every word.
If 60%+ of a batch fails the checklist, your prompt is probably too vague or the reference too weak. Go back and fix the prompt — don't just regenerate, that burns tokens on the same problem.
Common mistakes and weak prompts
The pattern I see people fail with, ranked by how often I've watched it happen:
Overloading the first prompt. Asking for scene, product, text, lighting, mood, brand colors, and three CTAs in one request. The model handles layered intent better than older models, but not infinitely. Start clean, refine with follow-ups. The cookbook's own guidance is small single-change edits, not everything in the opener.
Vague text instructions. "Include a tagline about the sale." You'll get a tagline — whichever one the model invents. Write the exact words.
Using a cluttered reference. If your product photo has four other objects in the background, the model might pull them into the output. Clean the reference or crop first.
Skipping the constraint line. Not telling the model what not to do is how you end up with extra hands, phantom products, and "AI-signature" backgrounds. Always include constraints.
Prompting for "cinematic" or "hyper-realistic" ad content. That pushes output toward movie-poster aesthetics, which for most TikTok Shop or affiliate ads is the wrong direction. You want platform-native, not stylized. Describe light and setting in plain terms.
Limits and trade-offs to know

Be clear about what this model won't do, so you plan around it and not into it.
It's not a video tool. You're making stills. If your final deliverable is a TikTok video, the output of this tutorial becomes frames in your CapCut timeline, not the finished asset.
Dense data breaks. Pricing tables with more than 5-6 rows, spreadsheet comparison grids, tiny numbers — the model still struggles. The fal.ai prompting guide flags ongoing limits on precise text placement and dense layout fidelity.
Latency is real. The gpt-image-2 model page notes complex prompts can take up to 2 minutes at high quality. Running 30 variants is an hour. Plan in batches.
Batch consistency drifts. Across 10 variants of the same product, expect small shifts in color, proportions, or label detail. Reference images help but don't eliminate it. Draft generator, not brand system.
No transparent backgrounds. Mentioned earlier but worth repeating. If your workflow needs PNG-with-alpha for downstream compositing, this model blocks you there.
Related Articles

Turn Product Images into Promo Clips with Omni Flash
How to use Gemini Omni Flash to turn product photos into short promo videos for TikTok Shop, Reels, and affiliate content.

Maya
Jul 8, 2026

Is GPT Image 2 Open Source?
Is GPT Image 2 open source? Learn what creators should verify about model access, official tools, API use, and safe alternatives in 2026.

Maya
Jul 7, 2026

Is GPT Image 2 Free? Pricing, Access, and Trial Options
Is GPT Image 2 free? Learn what to check about access, pricing, trials, ChatGPT plans, and creator use cases in 2026.

Maya
Jul 7, 2026

GPT Image 2 Speed: Is It Fast Enough for Creators?
GPT Image 2 speed matters if you make ad visuals, product images, or social variants. Here is how to think about speed in a creator workflow.

Maya
Jul 3, 2026

