Batch Your Whole Product Catalog Into Real Images
The Product Shots skill on Varosity generates product images and on-brand variations in bulk — from a product list or CSV — and returns hosted image URLs ready for your catalog, product pages, and ads. You can also hand it a real product photo and it restyles the scene while keeping the actual product intact, using the nano-banana model. One API key, synchronous calls, no job polling.
I kept running into the same wall when building out store demos and internal tools. You've got a CSV of 40 SKUs. Some have photos, some don't. You need images for every single one — white background for the catalog, lifestyle for ads, maybe a seasonal version. The old answer was either a Figma handoff to a designer, a photoshoot you couldn't afford, or wiring up three different image APIs with three different auth flows and hoping the outputs were consistent enough to matter.
What I actually wanted was: give me a list, give me back URLs. That's it.
What you're actually getting
For each product, you get a hosted image URL at cdn.varosity.ai — ready to drop into a database column, a product page, an ad creative. No downloading, no S3 buckets, no intermediate storage step. The URLs come back in the same call that generates the image. Synchronous. One round-trip per shot.
If you're running a 50-SKU catalog at one shot each, you're looking at maybe $0.50 to $4 in Varosity Credits total. Individual images run $0.01–0.08 depending on the model. Failed generations don't cost you anything.
The output for a CSV-based run comes back as a CSV too — one row per SKU, one column per variation URL. You pipe a spreadsheet in, you get a spreadsheet back with image URLs filled in. That's a thing you can actually hand to a non-technical person or load directly into Shopify.
The two modes — and when to use each
This is where most people make the wrong call early and then wonder why their product keeps coming out wrong.
If you have a real product photo — an actual image of your SKU — you use nano-banana with the referenceImageUrl parameter. That model preserves the actual product. The whole point of nano-banana is that your kettle stays your kettle. The scene changes. The background changes. The lighting shifts. But the object itself? Stays accurate. This is the workhorse for e-commerce. White background for the catalog page, warm kitchen lifestyle for the ad, marble countertop for the hero shot — all from one source photo.
If you don't have a product photo yet — maybe you're generating mockups before production, or you're building a demo store — you use gemini-3-pro-image or flux-1.1-pro and write a text description. gemini-3-pro-image is the clean, photoreal default. flux-1.1-pro skews toward richer, more stylized lifestyle scenes.
Don't use a text-prompt model when you have a real photo. I've seen this mistake a dozen times. You write a very careful description of your product, the model produces something that looks kind of like it, and then you spend twenty minutes iterating trying to get the logo placement right. Just pass the photo.
How the pipeline actually runs
The skill follows a clean three-stage loop.
First, you confirm the brief. Product list or CSV path, target style and background, aspect ratio (1:1 for feed, 4:5 for catalog, 16:9 for hero), and how many variations per product. Which items have a real photo to preserve. You confirm once, then the agent runs.
Second, for each product, it builds the request. Fresh generation means engineering a real prompt — product described precisely, scene, lighting, surface, angle. Keep the prompt clean. Busy scenes fall apart at thumbnail size. Restyle mode means the prompt describes the scene, not the product. Something like: "Place this product on a marble countertop, soft daylight, minimal props, 1:1." The product description lives in the referenceImageUrl, not the prompt text.
Third, it calls POST /api/v1/images — once per shot — and collects the URL from the response. No polling. The imageUrl is in the response body directly.
Here's what a fresh shot looks like:
POST https://varosity.ai/api/v1/images
{
"prompt": "Matte white ceramic pour-over mug, centered on a light-grey studio background, soft top-left softbox lighting, subtle shadow, e-commerce product shot, crisp focus.",
"model": "gemini-3-pro-image",
"aspect_ratio": "1:1"
}Response comes back with imageUrl — done.
And a restyle from a real photo:
POST https://varosity.ai/api/v1/images
{
"prompt": "Place this gooseneck kettle on a warm wooden kitchen counter, morning window light, a few coffee beans scattered, minimal lifestyle scene.",
"model": "nano-banana",
"aspect_ratio": "4:5",
"referenceImageUrl": "https://yourstore.com/skus/kettle.jpg"
}referenceImageUrl takes a public URL or a base64 data URI. PNG, JPG, WebP all work.
The request surface is intentionally minimal: prompt, model, aspect_ratio, and optionally referenceImageUrl. No width, height, seed, or steps knobs. Dimensions derive from the aspect ratio. If you need to lock a hero image across sessions, persist the returned URL — there's no seed, so re-running produces an equivalent but not bit-identical result.
The CSV flow in practice
Input looks like this:
sku,description,photo_url,variations mug-01,Matte white ceramic pour-over mug,,2 kettle-01,Gooseneck kettle,https://store.com/skus/kettle.jpg,1
Empty photo_url means fresh generation. Populated means restyle with nano-banana.
Output:
sku,image_1,image_2 mug-01,https://cdn.varosity.ai/img/mug-studio.png,https://cdn.varosity.ai/img/mug-lifestyle.png kettle-01,https://cdn.varosity.ai/img/kettle-lifestyle.png,
One row per SKU. One column per variation. Blank where fewer were requested. You load that into your store.
What to watch out for
The batch runs sequentially — not in parallel. There's a rate limit of roughly three requests per window, and if you hit it, you get a 429. The skill handles this by backing off and retrying. Don't try to parallelize the loop; you'll just get failures and have to retry manually anyway.
If you're generating shots where the on-pack text has to be legible — label copy, ingredient lists, product names on packaging — switch to ideogram-v3. The general image models will often garble small text. It's not a deal-breaker for lifestyle shots where the label is small in the frame, but if you're doing a close-up where the text is prominent, ideogram-v3 is built for that.
For variations to stay consistent across a SKU — say you want three different scene contexts for the same product — anchor them all to the same referenceImageUrl. That's what keeps the product consistent across the set.
Authentication needs the generate:image scope on your vsk_ key. Manage your keys at varosity.ai under API Keys. The scope lives on the key itself, so it works identically over REST or through the MCP server at https://varosity.ai/api/mcp. If you're running it through an agent via MCP, set the key once on the connection and it carries through.
Getting the skill
The skill lives in the Varosity public skills library. To install or update it manually:
curl -s https://varosity.ai/api/v1/skills/product-shots > ~/.varosity/skills/product-shots.md
It's set to auto_update: true, so if your runtime supports the refresh_skills MCP tool, call that at the start of a session and you'll always have the latest version. The source URL is https://varosity.ai/api/v1/skills/product-shots.
Once it's loaded, you call it with a plain instruction to your agent — something like: "Make product shots. Style: white-bg studio. Aspect: 1:1. Variations each: 2. Products: [your list or CSV path]." The skill handles the rest — brief confirmation, prompt engineering per product, API calls, URL collection, output table or CSV.
A 50-SKU run at one shot each takes a few minutes and costs less than a cup of coffee. Your catalog has images. You move on.