Back HexScope Lens

How AI Decides Which Brands It Remembers

POINT Key points
  • Exact AI recommendation lists almost never repeat, but appearance rate stays stable
  • Being remembered comes from a consistent category story across your sources

Type your own brand into ChatGPT, then ask the same question again ten minutes later. The list comes back reshuffled — new names, yours maybe gone. Do it a third time and it’s different again.

I’ve done this too, refreshing like the thing owed me an answer. And the honest reaction is: if it won’t hold still for one afternoon, does the AI even know my brand? Or am I just watching static?

But here’s what turned it around for me. The reshuffling is real, and it’s also hiding a second number that barely moves. One of those two numbers is the AI’s memory of you. The other is weather.

So let’s pull apart which is which — across three studies: one on brand recommendations, one on how models react to wording, and one on cloning actual people.

Across 2,961 Runs, One Number Held Still

The test nobody wants to run by hand, SparkToro and Gumshoe ran for us. Rand Fishkin’s team pointed 600 volunteers at three AIs — ChatGPT, Claude, and Google AI — and had them fire off brand-recommendation prompts.

The setup was broad on purpose. 12 prompt types, spanning both B2C and B2B (business-to-consumer and business-to-business), for 2,961 runs in total. Then they asked two very different things of the data:

  • Does the same list come back when you ask again?
  • And how often does a given brand show up at all, across everything?

So here’s what they found. The exact same recommendation list came back less than 1% of the time, and the ordering matched only about 0.1% of the time. Read that literally: track your rank, and you’re tracking something that reproduces once in a thousand tries.

But the appearance rate held. SparkToro calls it Visibility% — the share of runs where a brand shows up at all — and that number was statistically stable (meaning: measure it again and you get roughly the same figure, so it’s signal, not luck).

The numbers make it concrete. A high-visibility brand turned up in 97% of 71 runs. SaaS brands landed somewhere between 55% and 77%. Those aren’t die rolls; those are figures you can set beside last month’s, and the gap actually means something.

And there’s a detail I keep coming back to. They also ran 142 carefully varied prompts whose wording barely overlapped — a similarity score of 0.081, which is close to “these are basically different sentences.” The answers still converged onto a similar set of brands.

So the model wasn’t keying off the exact phrasing. It was keying off what the question was about. Which, for you, flips the whole exercise: “does AI remember us” isn’t a thing you read off one screen — it’s the appearance rate across a stack of same-intent runs, and that, unlike the ranking, sits still long enough to track.

It’s The Meaning Of The Question, Not The Wording

That last bit — wording barely mattered, meaning did — isn’t a fluke of one study. A separate paper, Errica and colleagues at NAACL 2025, went at it head-on.

They measured how a model reacts when you paraphrase a question, using two ideas they name plainly: sensitivity (does the answer change when you reword?) and consistency (does it stay right across rewordings?). It’s the same wobble SparkToro saw, isolated and put under a microscope.

The swing turned out small. Accuracy moved somewhere between 3.2% and 10% as they paraphrased. So phrasing does something — but not much.

So wording matters a little, and the category and meaning behind the question matter more. That lines up with the SparkToro convergence, and it’s oddly freeing. You don’t earn AI memory by guessing the one magic phrase a customer types. You earn it by being, clearly and repeatedly, the brand for this category — because that’s the layer the model is reading. Maybe I’m leaning hard on two studies, but they’re pointing the same direction.

Why “Seeming Like The Real Thing” Runs On Context, Not A Keyword

One more study, and I’ll flag up front that it’s an analogy, not proof. Park and colleagues at Stanford and Google DeepMind weren’t studying brands at all — they were cloning people.

They interviewed 1,052 people at length, then built an AI agent of each one straight from that transcript. Not a demographic label, but a whole qualitative interview per person — so the information density is nothing like a three-line persona.

Then they checked the clones against the . The agents reproduced each person’s own answers about 85% of the time — measured against that same person’s answers when they retook the survey two weeks later. People aren’t perfectly consistent with themselves either, so 85% is close to the human ceiling.

Now the caveat, because it carries weight: replicating a human’s attitudes is not the same task as an AI recommending a brand, and I don’t want to smuggle one into the other. But the mechanism underneath is suggestive. What made a clone “seem like the real thing” wasn’t one keyword or one label — it was the richness of the context feeding it. Thin the input and the resemblance thins with it.

Carry just that piece over. What makes an AI treat you as a real fixture of your category isn’t a single mention or a lucky keyword. It’s a consistent story about what you are, showing up across the many sources the model has read. So the move is to make your category role say the same thing everywhere — your site, the reviews, the roundups, the press — instead of a slightly different thing in each place.

Conclusion: Measure The Rate, Build The Consistency

Pull the three together and the picture’s pretty clean. The ranking is noise — it reproduces about one run in a thousand. The appearance rate is signal, stable enough to sit next to last month’s. And what moves that rate isn’t clever phrasing; it’s a consistent category story across the sources AI reads.

So two things, and only two.

Measure the right number. Stop screenshotting today’s rank. Pick a handful of same-intent questions, run each enough times, and watch the appearance rate — the share of runs you show up in — month over month. That’s the number that behaves like a memory.

Then feed that memory. Look at how your category role reads across your own site, third-party reviews, and the “best X for Y” roundups. If those sources describe you three different ways, the model has nothing steady to hold onto. Make them agree, and you’re writing the line the AI repeats back.

If you want the measurement side in more depth, the companion piece on why appearance rate beats rank walks through the same 2,961-run study from the metrics angle.


Sources

  • Rand Fishkin (SparkToro/Gumshoe), “NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility”, 2026-01-27, sparktoro.com
  • Federico Errica, Giuseppe Siracusano, Davide Sanvito, Roberto Bifulco, “What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering”, NAACL 2025, arXiv:2406.12334, arXiv
  • Joon Sung Park et al. (Stanford University, Google DeepMind), “Generative Agent Simulations of 1,000 People”, arXiv:2411.10109, 2024, arXiv
Share this article
Bluesky X
Back to all articles