“Can’t I Just Let ChatGPT Buy It?”
Buying anything online used to mean reading review after review written by people who are nothing like you. Lately I skip that and ask ChatGPT instead.
For the boring repeat purchases — toothpaste, detergent, cables — I’ve basically stopped forming opinions at all; I just do what the model says and click.
You’ve done a version of this too, right?
And the next step is already being built. Agentic commerce is a fancy way of saying you hand the AI your wallet and let it do the buying for you.
But here’s the bit I can’t shake. Does an agent with a budget shop the way a person does? If its taste is skewed in some AI-specific way, that matters a lot to whoever’s selling — and quietly, to whoever’s buying.
Somebody actually checked. So let’s look at what they found.
They Made The Models Go Shopping, Hundreds Of Times
The study comes from teams at Columbia University and Yale University, together with MyCustomAI, a company that builds measurement tools (R). They built a harness called ACES, ran the major models as real purchasing agents, and audited what those agents picked.
The setup is thorough. Six models from the ChatGPT, Claude and Gemini families; eight categories (fitness watches, toothpaste, washing machines, that sort of thing) with eight products in each; then hundreds of “which one do you buy” decisions, held up against how humans shop the same shelf.
Demand Piles Onto A Handful Of Staples
So here’s the first finding: agent choice is weirdly lopsided.
People disagree with each other, so taste splits and demand spreads out across a category — messily, but it spreads. The agents didn’t do that: they funneled demand onto a small handful of staples in each category and ignored almost everything else.
Eight products sitting on the shelf, and the model sees two or three.
If you sell things, that’s a scary sentence. Making the staple set or not is the difference between real AI-driven revenue and none at all.
And The Staple Set Reshuffles With Every Release
What’s worse is that this set doesn’t hold still; it gets rewritten every time the AI gets a version bump.
The study reports one Fitbit model picked 45% of the time under one version of Claude; it jumped to 77% under the next version, then sank to 6% under a different, newer model. Same product, same shelf. The only thing that moved was the AI.
Which means your presence can evaporate for reasons you had no part in and can’t negotiate with. I found that genuinely unsettling, and I think that’s the right reaction — it’s a thrilling little world we’re building here.
What Moves The Needle, And What Backfires
So is a seller just along for the ride? Not quite. The study also pulled out the levers that shift agent choice:
- Position matters even with no screen. In text-only exchanges, the order products came in changed how likely each was to get picked
- A “Sponsored” tag actively hurts. Products flagged as advertising were picked roughly 10 to 20% less often
- A platform’s own badge is powerful. An “Overall Pick”-style endorsement from the platform lifted selection sharply
- Rewriting the description moves share. When sellers reworked their product copy, the share of picks rose clearly across most of the models
That first one is worth sitting with; it means presentation order isn’t a “shelf layout” problem you get to wave away as UI.
Put the rest together and the shape is clear. Pushing harder with ads tends to work against you (like the salesperson who loses the room precisely by closing too hard), while tidying up your description and earning a legitimate third-party endorsement is the thing that actually lands.
Conclusion: The AI’s Shelf Rearranges Itself While You’re Not Looking
So AI shopping runs on logic that isn’t ours. It concentrates on a few products, reshuffles on model updates, and turns cold the moment you look like an advertiser — a difficult customer, this one. What do you do with that from the selling side? Two things, I think.
First, stop treating AI visibility as something you measure once. The study is pretty direct about this: which products get picked moves every time a model ships. So how the AI is handling your brand right now has to be tracked continuously, or you find out far too late.
Second, don’t aim the effort at the wrong thing. Reaching for ad tags to buy AI exposure made products less likely to be chosen, while writing a description a reader can actually parse, and stacking up honest third-party evaluation, is duller and works better.
Maybe the exact ACES numbers move as the models keep churning — that’s rather the point of the finding. But the direction holds: the shelf is short, and someone else keeps rearranging it. Knowing what this difficult new customer currently thinks of you is the cheap part. Start there.