All articles
ModelsJun 2, 2026·5 min read

How Vela chooses the right AI model for product visuals

A new AI model appears every few weeks. Here is how we decide which model makes which product image or video, and why you do not have to.

By Research Desk

Row of glass perfume bottles in different tints on a pale table, each lit differently as if in a test set-up

Short answer: there is no "best" AI model for product visuals. The right model depends on the job: your real product in a new setting, a mood shot, an image with a lot of text or a short video. At Vela we choose per job based on seven criteria, with label accuracy and cost per usable image at the top. You do not have to do anything: you describe what you need and the studio picks the model.

Every few weeks a new model appears that is supposedly better than everything before it. Pin your content production to one model from last year and it shows in the output. But blindly following every new model is just as risky. So here is how we choose.

Why the model matters

Two images from the same brief can look completely different depending on the model. Models differ mainly in:

  • How faithfully they reproduce your product, including the label and small print.
  • How photorealistic light, shadow and materials look.
  • How well they work with reference images, such as your packshot and brand imagery.
  • What they can do: images only, or video with smooth motion as well.
  • What a usable image costs, including the attempts you throw away.

Three kinds of models for product content

Model typeWhat it doesStrong for
Editing model with referencesTakes your product photo and changes the surroundings while keeping the productYour real product in new settings, variants, seasons
Text-to-image modelCreates an image from a descriptionMood, backgrounds, concepts without an existing product
Video modelTurns an image or description into short videoReels, Stories, short ads, product motion

For product visuals we almost always start from an editing model with your own product photo as the reference. Having a model "invent" a product from a text description is exactly how labels warp and details stop matching.

Our seven criteria

1. Product fidelity

The most important criterion. Do shape, proportions and colour stay exactly as on the real product? An image that shows your product differently from reality costs you returns and trust.

2. Labels and small print

Brands have text on their packaging. We test whether a model reproduces your label literally, with the right letters, spacing and capitalisation. General tools often guess here.

3. Photorealism and materials

Glass, metal, textiles and glossy packaging are hard. We check whether reflections, shadows and textures are believable or look more like a 3D render.

4. Working with references

Can the model combine several reference images, such as your packshot, another colourway and a brand image for the mood? That determines how consistent a series of images becomes.

5. Formats and resolution

Does the model support the ratios your channels need, from 9:16 to 4:1, and a resolution up to print quality?

6. Motion for video

For video, what matters is whether the product stays the same while it moves. A bottle that changes shape halfway through the clip is unusable.

7. Cost per usable image

What counts is not the price per generation but what an image you actually use costs, including retries and review time. A cheap model that needs five attempts is often more expensive than a pricier model that gets it right first time.

What does that look like in practice?

Model names change fast, so treat this as a snapshot from October 2026. Independent comparisons show exactly the trade-offs above. One comparison of FLUX.2 Pro and Nano Banana 2 for product shots found that Nano Banana 2 was stronger at editing with reference images and at reproducing label text correctly, while FLUX.2 Pro produced more realistic light and materials for shots without a reference (Any AI Studio, product shot comparison). For video, models such as ByteDance's Seedance are a category of their own.

The lesson is not that one model "wins". It is that every model has its own strengths, and the choice has to be made per job.

How we test a new model before it enters the studio

  1. We pick difficult products: lots of label text, glossy or transparent materials, fine detail.
  2. We run the same brief several times, so we do not decide on one lucky result.
  3. We check at full resolution: shape, colour, material, logo, label text.
  4. We calculate cost per usable image, including rejected attempts.
  5. Only if a model beats what we already have do we use it for the jobs it is strong at.

If you want to compare models yourself, use the same approach: test on your hardest product, not your easiest.

Why you do not have to choose a model

At Vela you describe what you need the way you would explain it to a colleague. The studio picks the model and settings that fit the format and channel. Your brand guardrails live in the studio, not in a single model. If we switch to a better model behind the scenes, your content stays recognisable. More on that in Brand consistency at generative scale.

That way you benefit from every improvement in AI without having to compare models every month yourself. And your content cycle stays fast, as described in Catching the wind: shipping content faster with AI.

See the difference with your own product

Curious how your product looks with the models we use today? Send us one product line and see how it works, or book an introduction.

Frequently asked questions

What is the best AI model for product photos?

+

There is no single best model. For your real product in a new setting, a model that can edit using reference images works best. For photorealistic mood shots without a reference, or for video, other models are stronger. The job decides the model.

Why doesn't Vela always use the newest model?

+

Newer is not automatically better for your product. We first test new models on real product jobs: does the label stay legible, are colour and proportions right, how many images are usable straight away? Only then does a model go into the studio.

Do I need to choose an AI model myself as a Vela client?

+

No. You describe what you need the way you would explain it to a colleague. The studio picks the model and settings that best fit the format and channel.

How do you test whether an AI model suits your product?

+

Take your hardest product, for example one with a lot of label text or a glossy material. Run the same brief several times on each model and check shape, colour, material, logo and label text at full resolution.

What happens to my brand if Vela switches models?

+

Your brand guardrails, such as products, colours and style, live in the studio rather than in one model. So your content stays recognisable, even when a better model is used behind the scenes.

Ready to ship content that catches the wind?

Book a live intro and see the studio in action.