← All articles

Does Rufus Read Your Listing Images? What's Documented and What Isn't

"Rufus reads your images" is now one of the most repeated claims in Amazon seller content, and one of the least sourced. It is repeated because it is plausible and because it sells services. This post does something the genre mostly avoids: it separates what Amazon has actually published from what practitioners have inferred from patents and testing, says which is which, and then gives advice that holds up either way — because most of what makes an image legible to a model is what makes it legible to a person in a hurry.

First, the name changed

On 13 May 2026 Amazon replaced the Rufus branding with Alexa for Shopping, folding the shopping assistant into Alexa+ and putting it in the main search bar rather than a separate chat panel. Amazon has said Rufus continues to operate behind the new front end, so the underlying system did not vanish; the label did. In Europe, where the assistant launched in beta in late October 2024 across the UK, Germany, France, Italy and Spain, which name you see depends on your marketplace and when you look.

The scale is Amazon's own reporting on earnings calls: roughly 250 million customers used Rufus during 2025, and by the second quarter of 2026 Amazon put usage at more than 350 million shoppers over the preceding twelve months, with interactions up several-fold year on year. Those are company figures, not independent measurement, but they are attributable — which is more than can be said for most numbers in this topic.

What Amazon actually documents

Amazon Science's own technical write-up of Rufus describes what it was trained and retrieves on: the product catalogue, customer reviews, community questions and answers, and public information from the web. It discusses a custom shopping LLM, retrieval-augmented generation and Trainium/Inferentia inference. It does not mention images, computer vision, OCR or multimodality. Neither does Amazon's original launch announcement, which uses the same four sources.

That is worth sitting with. As of August 2026, we have found no Amazon statement that the shopping assistant reads the pictures in your gallery. Amazon has also published no Rufus optimisation guidance for sellers at all — no help page, no checklist, no image spec. Any article that quotes Amazon on this is quoting something else.

What is observable, without needing a patent

The absence of a statement is not evidence of absence, and there is plenty of visible proof that Amazon runs vision models across listing imagery — just not proof about this particular consumer.

  • Amazon began replacing seller-written A+ Content image alt text with automatically generated descriptions, Europe first. Something at Amazon is looking at those images and describing them well enough to serve a screen reader.
  • Amazon Lens and Lens Live match a photograph taken by a shopper against the catalogue, and Amazon has stated the assistant is available inside that experience. Matching a photo to a product is computer vision by definition.
  • Amazon's A+ guidelines reject images whose text is unreadable on mobile — a rule that is only enforceable if something reads the text.

So: Amazon demonstrably applies vision to listing images. Whether the conversational assistant retrieves over the output of that vision when it answers "will this fit a 30-inch waist?" is the specific step nobody outside Amazon has confirmed.

What practitioners infer, and how well

The widely circulated version goes: the assistant uses OCR to pull text out of infographics and packaging, plus a visual-tagging layer to extract use context and attributes from photography, and treats that as citable content alongside your bullets. The evidence offered is Amazon patent filings and hands-on testing of listings.

Two cautions. Patents describe capabilities a company has claimed, not features it has shipped; large firms patent far more than they deploy. And listing-level testing struggles to isolate images from everything else on the page, because the assistant also has your title, bullets, attributes, reviews and Q&A — all of which usually say the same things your infographics say. A test showing the assistant "knew" a spec that appears only in an image is meaningful; most published tests are not that.

Our honest position: it is more likely than not that image content contributes, and it is not established. Treat it as a reason to write clearly on your images, not as a reason to redesign your set around a model whose behaviour nobody can specify.

The ecosystem has already priced it in

Tooling has moved regardless of what is proven. Helium 10 added a Rufus-oriented mode to its Listing Builder in early 2026 — its own write-up, dated 25 March 2026, describes a toggle that takes an ASIN, surfaces gaps in the questions shoppers ask, and reshapes copy to answer them. The same product ships Visual Product Intelligence, which reads an uploaded product image and auto-fills the listing fields from it.

That second feature is the more interesting evidence. A vendor is confident enough in image-to-attribute extraction to build a workflow on it. The capability is ordinary now. Assuming Amazon has it too is not a wild leap.

Making images machine-legible without making them ugly

Here is the useful part, and the reason none of the above needs settling. Every recommendation below improves the image for a distracted human on a phone, so the downside of being wrong about the model is zero.

  1. One claim per image, written as a sentence. "Fits standard car cup holders" beats "CUP HOLDER FRIENDLY!" for a reader and for anything parsing text. Fragments in all caps are worse on both counts.
  2. Numbers with units, spelled out. "1.9 L / 64 oz", "fits 76–96 cm waist", "IPX7 rated". Ambiguous numerals are where both humans and models guess.
  3. Real, high-contrast type at a real size. Sans-serif, dark on light, generous leading, not laid across a busy part of the photograph. Amazon's own A+ rule against text unreadable on mobile is the floor, not the target.
  4. Never let an image contradict the listing. If the image says 64 oz and the attribute field says 1.8 L, you have created a conflict that a person resolves by leaving and a model resolves unpredictably. Consistency across image, bullets and structured attributes is the single highest-value habit here.
  5. Fill the structured fields first. Material, intended use, compatibility and dimensions belong in Seller Central's attribute fields, where Amazon treats them as verified data. The image is the human-readable second copy, not the only copy. Our listing optimisation checklist covers those fields.
  6. Answer questions, not keywords. The shift the assistant represents is from matching strings to answering intents: who it is for, what it fits, how it compares, what it is made of. Your gallery should answer the six questions your reviews and Q&A show buyers actually asking.
  7. Do not stuff keywords into images. There is no evidence it works, it makes the image worse, and dense unreadable text is explicitly rejected in A+.

Note what is absent from that list: nothing about hidden text, nothing about invisible layers, nothing about writing for a machine at the reader's expense. Anyone selling you a "Rufus-optimised" look that a human finds harder to read is selling you a downgrade against an unproven upside.

What would change our mind

Concretely: an Amazon help page or Science post naming images as a retrieval source; or a controlled test where a fact appearing only in an image — not in the title, bullets, attributes, description, reviews or Q&A — comes back in an assistant answer, repeated across listings. Until then, write your images so a tired person on a train understands them in four seconds. That has always been the job. Graflio builds sets from the listing and its reviews and leaves every headline as an editable text layer, which mostly matters because it makes point four cheap to maintain. The claims themselves stay yours, and so do the rules in our guide to text on listing images.

Research your product, then generate the set

Paste one ASIN. Graflio reads the listing and its reviews, studies the best sellers around it, then renders a full set of on-brand infographics — every image still editable. €10 of credits free, no card.

Try Graflio free →