← All articles

A/B Testing Amazon Listing Images with Manage Your Experiments

Manage Your Experiments is the only way to A/B test a listing image on Amazon's own traffic, and it is free. It is also the tool most likely to hand you a confident-looking number that means nothing, because the thing that decides whether an image test works is not the design — it is whether the ASIN has enough sessions to separate two conversion rates from noise. Amazon publishes no traffic threshold, so this post is about what is documented, what has to be estimated, and how to read a result without fooling yourself.

Who can run one

Three conditions, all from Amazon's own material. You need a Professional selling plan. You need to be a brand owner enrolled in Amazon Brand Registry, internal to the brand and responsible for selling it in the Amazon store. And the ASIN must be eligible, which Amazon defines as belonging to your brand and having "received enough traffic in recent weeks to produce valid experiment results".

That third condition is the whole story, and Amazon deliberately does not quantify it. There is no published sessions-per-week number, no minimum order count, no category table. The eligibility check is simply whether the ASIN appears in the tool: open Seller Central, hover Brands, choose Manage Your Experiments, and start a new experiment. Products that cannot reach a conclusion in a reasonable window are not offered. Anyone quoting you a specific threshold has invented it or inferred it from their own catalogue.

Availability by marketplace is also worth checking rather than assuming. The tool started as a US feature and Amazon has extended experiment types to EU stores since; which types are live in which store has changed more than once. The Brands menu in your own Seller Central is the authority, not a blog post — including this one.

What the tool actually does

Shoppers arriving at your detail page are split randomly into two groups: one sees version A, one sees version B. Both are live listings with real money moving. You can test images — including the main image — as well as titles, bullet points, descriptions and A+ Content, and multi-attribute experiments let you change several elements at once. Results update once a week until the experiment ends, and it costs nothing.

Duration has two modes. The default runs "to significance": the experiment ends when Amazon judges there is enough data to declare a winner, which its material says can be as soon as four weeks. If you set a fixed length instead, Amazon recommends eight to ten weeks. Results report units sold, sales, conversion rate, units sold per unique visitor and sample size, plus the probability that one version is better than the other and a projected one-year impact for the winner.

What is worth testing

The constraint that governs everything is that each test costs weeks, so you get a handful per SKU per year. Spend them on differences big enough to move a conversion rate.

The main image, first and by a wide margin. It is the only image that affects both click-through from search and conversion on the page, so a change there compounds across the funnel. Angle, product scale within the frame, whether accessories are shown, and packaging versus bare product are all legitimate main-image variables inside the pure-white rules.

The second image. It is the first thing a shopper swipes to on mobile and the highest-leverage slot in the gallery. Testing what argument it makes — the core benefit, the size and fit question, or what is in the box — is a real test.

The order of the set. Moving your dimensions diagram from slot six to slot three is a cheap experiment that changes nothing about the artwork.

What is not worth a ten-week slot: a font change, a shade of blue, a headline reworded from "waterproof" to "water resistant", a slightly different drop shadow. These are real design decisions and the wrong instrument is being pointed at them. The effect sizes are far below what a normal listing can detect, so the result will be inconclusive, or worse, significant by accident.

The sample size reality

This is where most image tests die, and the arithmetic is not Amazon-specific. The number of sessions an A/B test needs rises roughly with the inverse square of the effect you are trying to detect: halving the improvement you want to catch multiplies the traffic requirement by about four. A test built to catch a 20% relative lift is a fundamentally different exercise from one built to catch 3%, and only one of them is realistic on a mid-volume ASIN.

Two practical consequences. First, if your listing gets a few hundred sessions a week, only large differences are detectable at all, which is another argument for testing whole concepts rather than details. Second, "no winner" is the most common honest outcome, and it does not mean the images are equivalent — it means the test could not tell them apart at this traffic level. Those are different findings and only one of them should change your design.

There is also a cost people forget: for the duration of the test, half of your real customers see the version that loses. On a high-traffic ASIN that is a rounding error. During peak season on your best seller, it is not, which is a reason to run image tests before Q4 rather than through it.

Reading a result honestly

What the tool showsWhat it means
Probability that B beats AA confidence statement, not proof. A 70% probability is a coin weighted slightly, not a decision.
Projected one-year impactAn extrapolation of a few weeks of one ASIN's data. Directional at best; do not put it in a forecast.
Conversion rate differenceThe metric that matters, but only alongside sample size. Read them together or not at all.
Weekly updatesAn invitation to stop early on a good week. Decline it.
No significant differenceThe test could not separate them. Not evidence they are the same.

Four discipline rules make the difference between a testing programme and a slot machine. Write down what you expect to happen and by how much, before you launch — a prediction you record is the only defence against explaining any outcome after the fact. Do not stop an experiment early because a mid-flight week looks good; peeking at a running test and stopping on a high point is the single most reliable way to manufacture a false winner. Change one substantive thing per experiment unless you are deliberately running a multi-attribute test, because a bundle of changes that wins tells you nothing about which change won. And check whether anything else moved during the window — a price change, a coupon, a competitor's stockout, a review spike, an ad budget shift. A test is only clean if the rest of the listing held still.

If your ASIN is not eligible

Most catalogues have more ineligible SKUs than eligible ones, and the answer is not to fake an experiment. Panel testing services put a design in front of a recruited audience and return preference data in hours — that is stated preference rather than purchase behaviour, so it is weaker evidence, but it is honest about being weaker and it works at any traffic level. The other option is a sequential before-and-after on a stable listing, which is genuinely confounded by season, price, rank and advertising, and should be treated as a signal to investigate rather than a result. And where the traffic simply is not there, apply what the category leaders have already validated: the top listings in your niche have run these tests, and their image sets are the visible output.

Research your product, then generate the set

Paste one ASIN. Graflio reads the listing and its reviews, studies the best sellers around it, then renders a full set of on-brand infographics — every image still editable. €10 of credits free, no card.

Try Graflio free →