VJOURNAL

BusinessGlobal DeskAugust 27, 2026

Infographics on a marketplace listing: the first frame sells the click, the rest answer objections

A listing's image stack is a script, not a gallery. The first frame wins the click in the grid; the rest remove doubts in the order they arrive. Text size, the line between picture and description, and rules that differ by platform.

Infographics on a marketplace listing: the first frame sells the click, the rest answer objections

Answer in brief

A listing's image stack is a script, not a gallery. The first frame wins the click in the grid; the rest remove doubts in the order they arrive. Text size, the line between picture and description, and rules that differ by platform.

3 sources
The first frame is judged at thumbnail size, so it can only carry one idea at a time.
Frames after the first should follow the order in which buyers raise doubts, not the order of your selling points.
Text baked into an image is not read by marketplace search, so every fact needs a text twin in the description.

Why a listing needs a whole image stack, not one picture

The short answer: the pictures on a listing are not decoration, they are a sequence of answers. The first frame exists to win the click inside a grid of near-identical squares. Everything after it exists to remove one more reason not to buy.

Nobody reads the description before opening the listing. A shopper sees a thumbnail, a price and a rating, and decides in well under a second whether your square is worth a tap. At that moment the picture is doing all of the selling on its own.

Once the listing is open, a second conversation begins. People swipe far faster than they read, and each swipe is a silent question: how big is it, what is it made of, what comes in the box, and what happens if it does not fit me.

So "make us some infographics" is not yet a brief. The brief starts with a written list of the objections your category actually produces, turns that list into frames in order, and only then opens a graphics editor.

The first frame has exactly one job

The first frame is not the place to explain the product. Its only job is to stop a thumb that is already in motion, in a row where twenty competitors are attempting precisely the same thing at precisely the same moment.

That means it has to show what the object is, what condition it is in, and what separates it from its neighbours in the grid. Size charts, materials and bundle contents belong further along; here they cost you the very click they were added to earn.

The usual failure is cramming the entire offer into frame one: a discount flash, a guarantee badge, three icons and an eight-word strapline. Shrunk to thumbnail size all of that collapses into grey texture, and the emptier listing beside yours takes the tap.

The test costs nothing at all. Scale the file down on your own screen until it is roughly the width of a fingernail. If the object stops reading before the words do, you need a different photograph rather than another layer of type over this one.

From the second frame on: objections in the order they arrive

The rest of the stack works like a script, and its running order should follow the order in which doubts appear in a buyer's head, not the order of your favourite selling points about the product.

In most categories scale comes first: something to compare it against, the object in a hand or in a room. Then material and finish, close enough to be judged. Then what is included. Then the conditions: care, compatibility, and what happens when the size turns out to be wrong.

You can source that order from evidence you already own. Pre-purchase questions in the chat and one-star reviews spell out, in the buyer's own words, exactly which question the images failed to answer in time.

A frame that answers nothing is not neutral filler. It pushes the answer somebody actually needed two swipes further away, and only a fraction of shoppers ever reach the end of the stack to find it.

Text that survives a thumbnail

The real constraint on a listing image is not the size of the canvas, it is the size that canvas will be viewed at. A caption set in a thin face at 24 pixels on a 1000-pixel artboard becomes a smear the moment it lands in a phone feed.

The working method therefore runs backwards: decide how many words can survive the shrink, then write to that number. Two or three words per frame, one idea, a heavy weight, and no hairline serifs anywhere near it.

Contrast matters more than point size. White type over a pale photograph is unreadable at any size, while the same words on a solid panel are perfectly fine. WCAG 2.2 asks for a 4.5:1 ratio for ordinary text and 3:1 for large text, and that is a sane yardstick here even though a marketplace listing is not formally a web page.

Then check it the way it will genuinely be seen: on a phone, outdoors, at a glance, rather than on a calibrated monitor in a darkened room. Most of your buyers are living in the first situation, not the second one.

What belongs on the image and what belongs in the description

Text baked into a JPEG does not enter the marketplace's own search. The platform's internal search reads the title, the attributes and the description, which are fields, not the pixels of your layout.

That gives you a clean split of duties. The image carries what has to be understood in a second; the description carries what has to be found, checked and copied. Materials, exact measurements, part numbers and compatibility all belong in text.

Duplication is fine and often useful. A size chart on a frame helps the eye, but the same numbers must also exist as text, or nobody can search them, quote them or paste them into a message asking you a question about them.

The comparison worth remembering is this: search engines do index the text inside a PDF, because inside a PDF text is still text. A raster image does not behave that way. A word on it is a picture of a word, and it cannot be your only copy of a fact.

Schema.org Product and where the image actually sits

If the same product also lives on your own site, it has a technical half: the structured data. Schema.org's Product type describes an item through fields such as name, description, offers and image.

The image property expects a URL or an ImageObject. It carries the picture itself, not the sentences printed across the picture. Claims about the product live in name, description and offers, where a machine can actually read them.

The practical consequence is blunt. A promise that exists only on an infographic does not exist for any system reading the page, although it exists very firmly for the buyer, with everything that follows from that gap between the two.

So every number you put on a frame should have a text twin: in the marketplace's description field, and in the structured data on your own site. Otherwise two versions of the same product quietly begin to drift apart from each other.

Platform rules differ, and the differences are not cosmetic

Aspect ratio, background requirements, whether text is permitted on the first frame at all, how many slots you are given: every platform sets these itself, and inside a single platform they shift from one category to another.

Check them in that platform's own seller documentation before the shoot rather than after it. Reshooting because the ratio was wrong costs far more than the ten minutes the requirements page would have taken to read properly.

Find out separately where the platform lays its own furniture over your picture: discount flashes, delivery badges, rating chips. They cover corners, so the composition needs margin that you can afford to lose without losing the subject.

One set of layouts for every platform almost never survives contact with reality. What works instead is building masters so that recropping to another ratio never cuts through the object that mattered in the first place.

Accessibility as a working quality filter

WCAG 2.2 was written for web content rather than for marketplace imagery, but its criteria make a useful checklist anyway. It is openly sceptical of text presented as a picture, and it is specific about contrast.

The images-of-text criterion asks you not to render as a picture what could remain real text. On a marketplace you often have no choice about that, which is exactly why a frame should never be allowed to become a wall of captions.

Many seller interfaces give you no alt field at all. Where that is the case, the description and the attributes are the only text-bearing carriers of meaning you have left, and they cannot be allowed to stay thin.

On your own site, always write the alt: short, factual, describing what is in the frame, with no keyword stuffing. It is simultaneously an accessibility requirement and ordinary common sense about how pages get read.

How a stack ends up looking cheap

Cheapness reads through inconsistency far more strongly than through budget. Different typefaces on adjacent frames, three shades of one supposedly brand colour, icons borrowed from three different families.

The second tell is ornament with no job to do: rays, sparkles, exclamation marks and frames that mean nothing but occupy the space where an actual answer to a buyer's question could have gone.

The third is superlatives with nothing behind them. "Best quality" and "number one" on a frame cannot be checked and do not persuade anybody, while a specific, concrete property in the same position does.

Consistency is cheaper than beauty. One grid, one typeface, two type sizes and a single panel style make a stack look deliberate even when the photography behind it is genuinely simple.

Model-generated frames

Generating listing images with a model looks like a saving, and it drags two different kinds of risk behind it: a factual one and a legal one. They are worth separating, because they are closed in very different ways.

The factual risk is that a model renders a plausible product rather than yours. A seam, a zip, a texture, the number of holes: any of them can drift, and a drifted detail here means returns and disputes rather than an aesthetic complaint.

The legal risk is quieter. A fully model-generated image may leave nobody holding an exclusive right to it, so when a competitor in your category lifts your frame there may be nothing solid to stand on.

The sensible middle is your own photography as the base with graphics layered over it: captions, callouts, dimension lines. That keeps the stack accurate and keeps the question of authorship straightforward to answer.

Telling whether the stack is working

A stack has two separate jobs and therefore two separate measurements. The first frame owns click-through from impressions; everything after it owns conversion from a viewed listing into an actual order.

Change them separately. Swap the first frame, the whole sequence and the title in one afternoon and no result afterwards can honestly be attributed to any single one of those three edits.

Look past the sales number as well. Pre-purchase questions and "did not fit" returns are a free report telling you which frame is missing or which frame is quietly misleading people about the product.

Give it the time the category needs to accumulate a comparable number of impressions. On a slow-moving product, comparing one week against the next is noise dressed up as a finding, and it will send you rebuilding the wrong frame.

What to hire out and what to keep

The Demand Audit is $50, runs 1-2 working days and includes one round on the findings document; access to your seller account is not included. It is the cheap way to learn which frames are missing before you spend anything on a shoot.

Listing Refresh is $70 over 2-3 working days with one round, and it does not include photography, ad spend or account management. The Marketplace + SEO System is $110 over 3-5 working days with two rounds on listings and copy, with photography and paid placement excluded.

When the calendar is tight there is Marketplace + SEO Express at $190 over 3 working days with one round, also without photography or ad spend. Managed Demand Growth is $200/mo on a monthly cycle with 30 days notice to stop, and advertising spend on the marketplace is not included.

Shooting the source frames stays a separate job: photography is explicitly excluded from Listing Refresh, the Marketplace + SEO System and Marketplace + SEO Express. Raise it in the brief, together with the list of frames you want, rather than after the layouts have already been built. The terms sit on /services/marketplaces.

Practical checklist

  • Write down the objections your category produces before you brief anyone on design.
  • Give the first frame one job and strip everything else off it.
  • Shrink each frame to thumbnail width and re-read it before you upload.
  • Move every measurement, material and part number into the description as text.
  • Open the platform's own seller documentation and confirm ratio, background and text rules before the shoot.
  • Change the first frame and the rest of the stack in separate rounds so results stay attributable.

Questions and answers

How many frames should a listing have?

As many as the category has real objections, and no more than the platform allows, since the slot limit is set by the platform and varies by category. Count questions rather than frames: each frame should close one question, and a frame without a question just pushes the next answer further away.

What does it cost to find out which frames are missing?

The Demand Audit is $50 and runs 1-2 working days, with one round on the findings document; access to your seller account is not included. The point of it is to get the list of missing frames before you have paid for a photo shoot.

Can I just put the materials and measurements on the image?

You can, and it often helps the eye, but it is not a substitute for the text fields. The platform's internal search reads the title, attributes and description rather than pixels, so the same numbers have to exist as text or nobody can find, copy or verify them.

Will one set of layouts work across every marketplace?

Almost never without adaptation. Aspect ratios, background requirements and whether text is allowed on the first frame differ between platforms and between categories, and each platform lays its own badges over your corners. Build masters with margin so recropping never cuts the subject.

Should I generate the frames with a model?

As a base, treat it carefully. A model renders a plausible product rather than yours, and a drifted seam or texture turns into returns. On top of that, a fully model-generated image may leave nobody holding an exclusive right to it, which makes copying harder to challenge.