“The AI Liked It. Does Anyone Else?”
More and more product copy starts life in a chat window. You paste the spec sheet, ask for three variants, pick one.
And then — this is the part I want to talk about — you ask the same model which variant is best.
I do this too. It’s fast, it’s free, and it feels like a second opinion.
But it isn’t really a second opinion, is it? It’s the same reader grading its own homework. And if every brand in your category runs that loop (which, judging by my inbox, they do), the whole category slowly converges on one taste.
Someone actually measured that taste.
The 89% That Started This
The study is “AI-AI bias: Large language models favor communications generated by large language models”, published in PNAS (the Proceedings of the National Academy of Sciences — the US national academy’s journal, which runs everything from the natural sciences to the social sciences). The setup is simple, so let’s walk it. We take the same thing described twice — once by a person, once by a model — then ask an LLM (large language model, ie the thing behind ChatGPT) to pick the better version. They ran it across three kinds of text: consumer product descriptions, academic paper abstracts, and movie plot summaries.
So here’s what came out. GPT-4 chose the GPT-4-written version 89% of the time on consumer product descriptions, 78% on academic papers, and 70% on movie plots.
Then comes the useful part. They put the same pairs in front of 13 human research assistants, who went for the AI-written version 36% of the time on consumer goods, 61% on academic papers, and 58% on movies.
Look at the consumer-product number.
The model says 89, the people say 36. That’s not a rounding difference; that’s two different tastes wearing the same word, “good”. And it’s the number that matters most here, because product copy is the thing marketers actually ship.
What A Model’s Score Is Actually Made Of
AI-AI bias is basically this: models lean toward text an AI wrote — their own output included — over text a person wrote. Which, for you, means a high score can partly reward shape — the phrasing and structure models find easy to read — rather than anything a customer would notice.
But it doesn’t stop at style. The same tilt shows up at the level of brands.
A 2024 paper at EMNLP (one of the main natural-language-processing research conferences), “Global is Good, Local is Bad?”, tested brand recommendations across GPT-4o, Llama-3-8B, Gemma-7B and Mistral-7B. Every model tied global brands to positive attributes disproportionately. And the recommendations split along income lines: luxury brands went to high-income countries 88-100% of the time, non-luxury brands to low-income countries 84-98%.
Those are specific experimental conditions (and I wouldn’t assume every AI product behaves the same way). Still, the direction is worth sitting with. Because it means that when you ask a model to grade brand copy, you’re not getting a clean read on the copy. You’re getting the copy plus whatever the model already assumes about your brand, your category, and the market you’re selling into.
Which is a strange thing to hand your creative decisions to.
Why Agents Make This More Urgent, Not Less
Harvard Business Review’s “AI Is Upending Marketing on Two Fronts” looks at the same shift from the marketer’s side. Two things are happening at once: conversational AI is taking over work that search and websites used to do, and AI agents are moving toward making the purchase decisions themselves. The piece cites research showing online search dropped by about 20% after ChatGPT arrived, and argues the job is moving from SEO to GEO (generative engine optimization — optimizing for the answer a model gives, not the list of links Google shows).
But there’s a second split coming, and it’s the awkward one.
The “customer” and the “consumer” may stop being the same entity. If an agent picks the product and a person uses it, you’re writing for a reader who buys and a reader who lives with the thing — and, per the study above, those two readers don’t share a taste.
So the tempting conclusion is “fine, just write for the model.” That’s exactly what the PNAS numbers argue against. A 53-point gap between what a model rewards and what people pick isn’t a rounding error you can optimize your way through.
Model scores are a useful signal. They’re just not the verdict.
Conclusion: Keep Two Scoreboards, Not One Score
Being visible to the models is a real, buildable asset — the search shift and the agents both point that way, and it’s something you can go check today. What you can’t do is fold it into the same number as human trust, because on the product descriptions the model’s pick and the humans’ pick went opposite ways, 89% against 36%.
So: run the same product description past ChatGPT, Gemini and Perplexity, and track how often you get quoted or recommended. That’s scoreboard one, and my hunch is that most teams (mine included, for a long stretch) have never sat down and looked at it.
Keep scoreboard two where it already lives — ad A/B tests, customer interviews, the questions people ask before they buy. If you let scoreboard one absorb it, you’ll drift toward copy that reads beautifully to a model and thinly to everyone else.
Maybe I’m wrong about how durable the 89% is; it’s one study, on one model generation. But the cheap move is to keep the two measurements apart, and I don’t see much downside in doing it now.
Sources
- Laurito, W. et al., “AI-AI bias: Large language models favor communications generated by large language models”, Proceedings of the National Academy of Sciences, Vol. 122(31), 2025, pnas.org
- Kamruzzaman, M., Nguyen, H.M. & Kim, G.L., “‘Global is Good, Local is Bad?’: Understanding Brand Bias in LLMs”, EMNLP 2024, aclanthology.org
- Puntoni, S., “AI Is Upending Marketing on Two Fronts”, Harvard Business Review, 2026-02-23, hbr.org