Back HexScope Lens

One Product-Copy Rewrite Moved The AI's Picks Up To 14.9 Points

POINT Key points
  • An AI shopper turns away from anything tagged 'Sponsored'
  • Fixing the words beats paying for the slot

“Apparently The AI Does The Shopping Now”

Somebody says it in a meeting, everyone nods, and the topic moves on.

Then a week later somebody asks the follow-up — fine, so what do we actually change on the shelf? — and the room goes quiet. I’ve been in that silence. I didn’t have an answer either.

All the instincts we built for humans (get the reviews up, get the price display right, buy the placement) had never been tested on a buyer that reads a JSON list and doesn’t have eyes. Nobody had checked whether any of it transfers.

I once handed a model a headline I was quite proud of, written to make a person feel something, and watched it get ignored completely.

Somebody built a fake marketplace and measured the whole thing properly, so let’s go look at what came out.

What Does An AI Shopping Agent Actually Go On?

Placement, sponsorship labels, and the platform’s own endorsement badge — and it reacts hard to all three.

In the simulated marketplace, which the team ran in 2025, a “Sponsored” tag pushed the probability of being picked down, while a platform’s own “top pick” style badge pushed it sharply up (Allouah, Besbes et al., arXiv:2508.02630). The authors are split across MyCustomAI, Columbia University Graduate School of Business and Yale University.

So some of what human advertising taught us runs backwards here. The signal that says “someone paid to put this in front of you” is exactly the signal an agent discounts.

Position Bias Doesn’t Respond To Human Ad Sense

The experiment ran on a simulated marketplace the authors built called ACES, where placement, tags, price and ratings could all be shuffled at random while they measured what that did to each product’s chance of being selected. Frontier models did the buying — GPT-4.1, GPT-5.1, Gemini and Claude among them.

Position bias turned out to be strong, and the direction of it differs by provider and even by model version. Between GPT-4.1 and GPT-5.1 the preferred slot nearly reversed.

Strip the images out entirely and hand the model a plain JSON list — the “headless” condition — and the same bias shows up again. That’s a property of the model, then, not of how it reads a page layout.

And telling it outright to ignore position barely helped. There’s a ceiling on what you can fix with prompt wording, which is worth knowing before someone on your team proposes fixing it with prompt wording.

One Rewrite, And The Share Of Picks Moved

Here’s the experiment I keep coming back to. The authors let an AI sales agent read the competitors’ sales data, then rewrite its own product description exactly once, and they measured the effect against six different AI buyers.

Buying modelGain in share of picks
Gemini 3 Pro Preview+0.32pt
Claude Sonnet 4+3.66pt
Claude Opus 4.5+7.38pt
GPT-4.1+8.37pt
Gemini 2.5 Flash+14.79pt
GPT-5.1+14.89pt

Five of the six moved by a statistically significant margin. Across all the experiments they ran, 33% produced a significant increase off that single rewrite.

Which is not “always works”. But a lever that pays off a third of the time, for the cost of rewriting a paragraph, is about as cheap as marketing levers get.

What The Numbers Don’t Cover

Four limits, and I’d say all four before this table goes into anyone’s deck.

It’s a simulated marketplace. The effect size on a real storefront, with real inventory and real buyers, could sit somewhere else entirely.

Only 33% of the experiments reached significance, so the majority showed no visible movement at all.

The position bias flips direction between model generations, which means the slot that works today has no obligation to work next month (about as durable as reading your horoscope once and planning the decade around it).

And two of the five authors work at MyCustomAI, which sells e-commerce optimization for AI agents. “Rewrite your descriptions and your share goes up” is the conclusion that sells that product. Take the finding; discount the gradient it sits on.

Rewrite Your Descriptions For A Reader That Has No Eyes

Start with the copy on your main products. Read it once as a model would: is the price there, are the specs there, are the questions a buyer would ask in a comparison actually answered in plain words? Buying placement and sponsored slots is the older reflex, and this data says it carries less weight than it used to.

Be careful specifically with the sponsorship label. The tag itself dragged selection probability down in this experiment, so paying for the slot without also fixing the words underneath can buy you a worse outcome than doing nothing.

And since position bias moves every time a model ships, treat this as something you check on a schedule across several assistants, not a job you finish once. That’s how you notice the shelf being rearranged while it’s happening rather than a quarter later.


Sources

  • Amine Allouah (MyCustomAI), Omar Besbes (Columbia University Graduate School of Business), Josué D. Figueroa (MyCustomAI), Yash Kanoria (Columbia University Graduate School of Business), Akshit Kumar (Yale University), “What Is Your AI Agent Buying? Evaluation, Biases, Model Dependence, & Emerging Implications for Agentic E-Commerce”, arXiv:2508.02630, 2025-12-17, arxiv.org (Experiments run on ACES, a simulated marketplace built by the authors, with placement, tags, price and ratings randomized. A “Sponsored” tag lowered selection probability while a platform endorsement badge raised it sharply. Position bias was strong and its direction differed by provider and model version — nearly reversing between GPT-4.1 and GPT-5.1 — reproduced in a headless JSON-only condition and barely reduced by instructing the model to ignore position. Letting an AI sales agent rewrite the product description once produced statistically significant share gains in five of six buying models: Gemini 3 Pro Preview +0.32pt, Claude Sonnet 4 +3.66pt, Claude Opus 4.5 +7.38pt, GPT-4.1 +8.37pt, Gemini 2.5 Flash +14.79pt, GPT-5.1 +14.89pt; 33% of all experiments showed a significant increase. Simulated environment, so effect sizes may differ from real e-commerce; two of the five authors are employed by MyCustomAI, which sells e-commerce optimization for AI agents.)
Share this article
Bluesky X
Back to all articles