Aims
All posts
Production8 min read

Why AI Product Images Get Your Product Slightly Wrong

It is not a prompt problem. It is a description problem.

Henry Sedgwick

Henry SedgwickProduct marketing

A sage green cosmetic bottle with a brass cap on a limestone plinth in front of a mirror, where the reflected bottle has a silver cap, a sloping shoulder and different proportions

If your AI product images keep coming back with the wrong product in them, you are not prompting badly. The bottle is a little too tall. The cap is the wrong metal. The label looks like your typeface from across a room and nothing like it up close. The reason is worth being blunt about: unless you hand the model your actual photograph, it is not copying your product. It is drawing something that matches your description of it.

Those are different jobs. A description is a lossy summary, and everything you forgot to mention gets invented, confidently and attractively. Words control the scene, but only pixels control the product, and if a detail matters (the wordmark, the closure, the stitching, the net weight line) it has to arrive as an image. So the fix has two halves: lock the product with a real reference image instead of adjectives, and take the scene from a photograph a person actually lit. I ran the test below to show what each half is worth.

01

I ran the test on a real product

Loud Bite make popped lentil chips. Their packs are a good test because they are busy: heavy black wordmark, scribbled doodles, starburst badge, flavour name, protein banner, weight mark. I wrote the most careful description of their real photograph I could manage (about 250 words) and generated from that alone. Then I generated again, same scene brief, with the photograph attached.

Same product, three ways

ControlThe real photograph
Five Loud Bite popped lentil chip packs standing in a row on white, in green, red, pink, turquoise and yellow, each with a heavy black wordmark and a flavour name
The actual product. Note the chunky slab wordmark and the round, disc shaped chips.
DescribedWords only, no reference
Three generated snack packs in pink, turquoise and yellow that resemble the real product but use a brush script wordmark and show triangular tortilla style chips
A handsome bag. Somebody else's. Brush script wordmark, triangular chips, no seal band.
ReferencedSame brief, photo attached
Three generated snack packs matching the real product closely, with the correct heavy slab wordmark, zigzag doodles, round chips and contrasting coloured top seal bands
The product survives. Wordmark, doodles, seal bands, chips and the plus sign all carry across.
Fig. 1 Left: the real photograph. Middle: generated from a careful 250 word description, no reference. Right: same scene brief with the photograph attached. Middle and right were made minutes apart with the same model.

The middle image is the important one, and not because it is bad. It is good. I would happily put it in a grid. It is also, in twelve countable ways, not the product Loud Bite sell. That is the trap. A broken image gets deleted. A plausible one gets shipped.

Twelve things the description could not carry

The generated pack shot annotated with numbered pins marking the wordmark, the flavour line, the chips and the pack proportions
1
2
3
4
5
6
1
The wordmark is a different typeface. The real one is a heavy geometric slab with a spike cut into the last letter. This is painted brush script. The most valuable thing the brand owns, and it did not survive.
2
The chips are the wrong shape. Loud Bite sell round popped discs. These are triangles. Anyone buying from this image opens a bag of something else.
3
The flavour reads "sea salt and malt vinegar". The pack uses a plus sign. I wrote "and", so the model was obedient and wrong, which is the honest shape of this failure.
4
The doodles changed species. Zigzag squiggles on the real pack, lightning bolts here. I said "scribbled doodles" and got somebody else's idea of one.
5
No contrasting seal band. Every real pack has a purple or navy strip across the top seal. I never mentioned it, so it does not exist.
6
The proportions drifted. Real packs are wide and squat. These are tall and narrow. Nobody writes an aspect ratio into a prompt.
Six more unpinned: the starburst flare behind the flavour name is gone, so are the chevrons either side of 100G, the protein banner is the wrong colour on two packs, "BIG CRUNCH" lost its yellow second word, and there is no barcode or recycling mark. Twelve, from a description I genuinely tried hard on.
Fig. 2 Each of these was either absent from my description or too fiddly to put into words. The model filled every gap with something reasonable, and reasonable is the problem.
"We shipped a month of paid social before someone in the warehouse pointed out the pump on our bottle was the wrong shape in every single ad. Nobody in marketing had noticed. They handle those bottles every day." Kofi, haircare, six people

02

Why AI product images drift in the first place

Knowing what happens tells you which fixes are real and which are superstition. When you generate from text, nothing about your product is in the system. The model reads your words, finds the region of everything it has seen that best matches them, and draws from there. Your brand is not in that region. Ten thousand other snack bags are.

Two different operations hiding behind one text box

Generate from text

Edit from an image

What the model receivesYour sentences onlyYour actual pixels
Where the product comes fromEverything it has seenThe photo you attached
Small print and logosInvented every timeMostly preserved
Good forConcepts and moodboardsAnything a customer buys from
Fig. 3 Most tools put both behind the same box, so plenty of people are using the wrong one without knowing there was a choice.

OpenAI ships a fidelity dial

Its image editing API has a high input fidelity setting, documented as being for "editing images with faces, logos, or any other details that require high fidelity". The existence of the setting is the admission. The docs also note only your first attached image keeps extra texture richness, so order matters.

Google labels its own tool experimental

Merchant Center Product Studio carries a live warning that it "may generate unexpected images and videos and works best with certain products", and it will not process images containing hands.

Reference slots are capped

Gemini's image models limit how many references you can attach, six object images on the Pro tier at the time of writing. Worth knowing before you plan a workflow around feeding it twenty angles.

None of this is in the marketing. All of it is in the documentation. These are not broken tools, they are tools telling you where their edges are.

29.0%

Best product fidelity pass rate of four leading edit models, Photoroom benchmark, July 2026

850

Real products tested, across clothing, footwear, bags, jewellery and accessories

3,400

Generations, each judged by three or more trained annotators, one flag and it fails

A broken image gets deleted. A beautiful, plausible, wrong one gets shipped, printed and paid for.

03

What people actually report going wrong

Almost every article on this subject says "users report" and then names nobody. Here are three threads I read in full, with dates. A product designer on r/productdesign summed the whole category up in one line in November 2025: every generator they tried distorted small fonts on packaging.

The SKU drifts across a batch

An operator writing up 100+ ecommerce generations on r/StableDiffusion in July 2026 found newer models better at everything except keeping one pill case identical across fifty shots. The handle changes slightly, the logo warps. They went back to older models purely for predictability.

A reference image is not a force field

On r/generativeAI in April 2026, someone generating a serum bottle at 2K with the real photo attached still could not hold the label. Parts came out blurry, and in places the model invented text. That is the honest limit of the fix, and why small print wants compositing rather than generating.

Reflections nobody asked for

The same operator flagged glass, metal and glossy packaging as worst, because the model keeps adding reflections of windows and lamps that were never in the room. Foil pack or glass bottle, budget extra checking time.

Low-traffic threads, all of them, which is rather the point. Not a scandal, just a daily annoyance people quietly work around.

04

Getting the product right is only half the job

Lock the product perfectly and you have solved accuracy, not quality. You now have a correct bottle floating on a grey gradient, which is the house style of AI output and the reason people spot it instantly. A reference photo fixes the product. Only a person ever made the scene.

Three glossy skincare bottles in orange, pale blue and yellow tumbling diagonally across a gradient background from sky blue to cream, covered in fine condensation with soft coloured caps
Fig. 4 Studio LIIT for Glowery, pulled from the AIMS library. Look at the decisions: bottles tumbling rather than standing, cropped so hard that two run off the frame, condensation as the whole texture story, a background shifting cold to warm across the diagonal. Ask a model for "skincare photography, bright and fresh" and it centres three bottles and lights them evenly, because that is the average of everything it has seen. This is not the average of anything.

05

Count what has to survive

A KOR Immune Plus juice shot bottle standing beside its matching carton on a pink surface against a pale blue background, both showing the KOR wordmark, the ingredient lines and a USDA Organic seal
Fig. 5 KOR Shots, from the AIMS library. Count the load bearing details in one simple two object shot: the reversed K and R in the wordmark, the USDA Organic seal, both ingredient lines, the 1.7 fl oz on the bottle against 12 x 1.7 fl oz on the box, the teal and orange split repeating across both. Any of those coming out wrong is a listing problem, not a taste problem.

06

How to actually lock it down

  1. 1

    Start from a photograph, not a paragraph

    One clean, sharp, well lit shot of the real product on a plain background. The highest leverage asset you own, and it costs a phone and a window.

  2. 2

    Attach it, and put it first

    Use the edit or reference path rather than the text box. Turn any fidelity setting up. Where order matters, and it usually does, your product goes in slot one.

  3. 3

    Give it more than one angle

    Front, three quarter, and a close crop of the label. Ambiguity is what gets filled in with invention, and one flat angle leaves a lot ambiguous.

  4. 4

    Ask for a scene, not a redesign

    Your words should describe surface, light, angle and mood. The moment a sentence describes the packaging itself, you have invited the model to redraw it. My reference prompt said "do not redesign, redraw or reword any part of the packaging" and that line earned its place.

  5. 5

    Run the flick test before it ships

    Put the generated image and the real photograph side by side and flick between them. Drift you cannot see in isolation is obvious in a flick. Check the wordmark, any seal, the net weight, the shape of the product itself.

Nothing exotic here. The first two steps alone took me from twelve visible errors to a handful of soft ones in the small print.
  • Composite when the stakes are high: keep the real product layer untouched and let the model build only the background around it.
  • Never let a generated image be the main marketplace listing photo, where accuracy rules are strictest and enforcement is automated.
  • Keep one approved reference per SKU in a fixed place, so the team generates from the same file rather than whatever screenshot is nearest, and re-run the flick test whenever you change models.
"The thing that fixed it for us was boring. We reshot one hero image of each SKU properly on a windowsill and put them in a shared folder. The AI stuff went from unusable to fine. Same tools, same week." Ilse, ceramics, solo

07

Where AIMS fits

Short, because the method above works with whatever you already have. AIMS exists because doing both halves properly is fiddly and nobody wants a checklist as a job. A product reference attaches to your brand once and rides along with every generation after it, so the SKU stops drifting without anyone remembering to stop it.

The other half is why I built it this way. Discover is a free library of real work (the Glowery shot and the KOR listing above both came out of it) and when you pick one, the analysis step writes down what the photographer decided: crop, angle, light, surface, colour logic. That analysis is what reaches the model, next to your locked product. A scene a human composed, on the product you actually sell.

08

Common questions

Why does AI change my product when I described it so carefully?

The care goes into the wrong channel. A description compresses your product and generation decompresses it using a different reference set than yours. My 250 word attempt still lost twelve things, and some of them, like the exact proportions of a bag, are not really writable. Attach the photograph instead.

Will a better prompt fix product accuracy?

Only up to a point, and the point is low. Prompts control scene, light and framing well, but they are a poor instrument for specifying an object that already exists. Reach for a reference image first, then spend prompt effort on the scene.

Why does the text on my packaging come out garbled?

Small text is where drift is most visible, because you can read it. Vendors document the limits directly, including underperformance on non-Latin alphabets and on text whose colour or style varies. Keep real type as a layer you composite on top rather than asking a model to render your legal copy.

Why is my product a slightly different colour in AI photos?

Two causes worth separating. The model may have a house tint it applies on every pass, which stacks if you run an image through twice. And it may substitute its own idea of a colour name for your shade, because "sage green" is a region rather than a value. Attach the reference, avoid re-editing an edited file, and check against the real thing before it reaches a listing.