Back HexScope Lens

One Line About The Buyer Changes Which Brands The AI Recommends

POINT Key points
  • Naming the buyer swaps up to 75% of the mid-market shortlist
  • 2,000 runs across 10 personas, 8 prompts and 3 model setups

“Whose Numbers Are These, Exactly?”

You pulled the number yourself, and you still can’t defend it the moment somebody pushes on it.

What share of prompts named you: that you can show. How much it moved since last month: that too. Two real numbers.

But push once on whose question produced the line and the ground goes soft.

The assumption held for me for years — identical prompt text meant identical conditions. Same words in, same playing field.

One line about the asker went on the front of the question, and a different set of brands came back. Two thousand runs of it, and the middle of the market is where the names moved.

Does The AI Recommend Different Brands To Different People?

It does. Keep the question word for word, add “you are a UK small-business owner” in front of it, and the overlap between the two brand lists falls by 0.12 to 0.20.

Those lists moved most in the middle of the market. Mid-market brands had up to 75% of their names swapped out between personas, while category leaders kept roughly 80% of their spots whoever was asking.

The work is a preprint — posted before anyone outside has reviewed it — from Will Jack, Noah Lehman, Keller Maloney and Sarah Xu at Unusual.ai, put on arXiv on May 28, 2026 (the paper).

What they pointed it at was commercial chat that answers off live search results. Ten personas × eight prompts × three model setups × ten repetitions, for 2,000 runs in total.

A persona here is one line of setup: the person asking is this kind of person. They hand-built ten of them out of industry, company size, job title and region — a US solo founder, an enterprise procurement lead, a UK small-business owner, and seven more like that.

The Top Seat Doesn’t Move. The Middle Has Vacancies.

Prominence turned out to be what decides how much the asker matters. The paper sorts brands into tiers by how well known they are, then reports how much of each tier’s lineup got replaced when the persona changed.

Across the three model setups, the ranges came out like this:

  • Category leaders: 20-29%
  • Mid-market: 39-75%
  • Regional players: 27-33%

The leader gets named for everybody, so that seat is nailed down.

And the middle is where seats empty out and refill as the asker changes.

But that mostly isn’t bad news. A seat that nobody has nailed down is a seat you can take.

How Much Do You Get Back If You Just Ask Twice?

The overlap measure here is the Jaccard index, which scores how much two lineups have in common on a 0-to-1 scale — the closer to 1, the more the same brands are standing there.

Across personas, that overlap ran 0.22 to 0.35. If you hold the persona fixed and just repeat the same prompt, it still only reached 0.42-0.51.

So the same person asking the same thing gets back a lineup that’s close to half different each time. The persona effect stacks on top of that.

Which makes “we were in the recommended set” off a single run about as informative as one round of rock-paper-scissors. If you aren’t the category leader, dropping in and out of that set month to month might be the structure of the thing rather than noise in your measurement.

I think the order matters here: repeats first, personas second. Miss either one and the gap you’re looking at can’t be read.

How Far Does Any Of This Travel?

Every author on the paper works at Unusual.ai, a vendor in AI brand visibility, and “you have to measure per persona” is a conclusion that happens to suit them. Five more things bound how far these numbers reach.

  • Peer-review status isn’t stated anywhere — it’s a preprint
  • Ten hand-built personas, English only, concentrated on the US, the UK and the EU
  • One day of measurement, so drift over time goes untested
  • The long-tail specialist tier was underpowered in every condition (fewer than 30 cases per cell), so it reports no number at all
  • The Anthropic setup covers only 4 of the 8 prompts

And none of it is causal. What you can say is that changing the asker changed the output. “Write a persona and you’ll win the recommendation” is never on offer.

The two providers didn’t produce the same recommendations, either. The Anthropic model shifted further with the asker (the 0.20 drop), and 43-52% of its recommendations weren’t tied back to anything retrieved. The OpenAI setups run 8-29% on that same measure.

The authors’ guess is that it comes down to how much of an answer leans on knowledge from training rather than from search, and they say plainly that the mechanism is unsettled. That reads to me as a reason to measure each model on its own terms. But a single day of runs is thin ground, and maybe I’m leaning on it harder than it can hold.

Ask The Same Question As Two Or Three Of Your Customers

Held as one line on a chart, AI visibility flattens all of that movement into an average. And for a brand in the middle of its market, the average is sitting on top of a lot of turnover.

Getting under that average costs one change in how you send the prompt. Write down two or three of your core customer types, send your usual question with each of them as the opening line, and lay the results side by side.

Since averaging cancels the tier where seats are open against the tier where they’re taken, keep the three apart.

And that’s when you can finally see which customer’s answer your name is missing from.

Whose question produced this line? For my money, that belongs written on the chart, right next to the number.


Sources

  • Will Jack, Noah Lehman, Keller Maloney, Sarah Xu (Unusual.ai), “Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit”, arXiv preprint 2605.30207v1, May 28, 2026, arxiv.org (factorial design of 10 personas × 8 questions × 3 model configurations × 10 repetitions = 2,000 runs. The configurations are GPT-5.4-mini at low and at high reasoning effort, plus Claude-sonnet-4.6 at low reasoning effort. Adding the persona opener drops the Jaccard similarity of the recommended set by −0.12 [95% interval −0.153, −0.091], −0.16 [−0.194, −0.139] and −0.20 [−0.263, −0.156], and none of the three intervals contains zero. Similarity across different personas runs 0.22 to 0.35; within the same persona, 0.42 to 0.51. Recommendation turnover by prominence tier: category leaders 23%/29%/20%, established challengers 5%/13%/23%, mid-market 75%/67%/39%, regional players 27%/33%/—. In the abstract’s own wording, category leaders hold about 80% consistency while mid-market brands turn over up to 75%. Recommendations not tied to a retrieved source run 43% to 52% on Anthropic and 8% to 29% on OpenAI. Limits: every author works at Unusual.ai, an AI visibility vendor with a commercial interest in this conclusion; the peer review status is not stated; the long-tail tier is under-sampled in every cell, at fewer than 30 brand events per cell; the sonnet cells cover only 4 of the 8 prompts; the personas are 10 hand-written ones, in English only, concentrated in the US, the UK and the EU; variations in the wording of the opener are untested; the measurement covers a single day, so drift over time is unevaluated; and the mechanism the authors suggest — that the persona acts on generation from prior knowledge rather than on retrieval — is unconfirmed)
Share this article
Bluesky X
Back to all articles