Back HexScope Lens

ChatGPT Names A Different Brand Every Time. Measure Anyway.

POINT Key points
  • The recommendations really do shuffle, but part of that shuffle is your measuring method
  • What holds steady is appearance rate across repeats, not the single-shot ranking

You ask ChatGPT “what are the best tools for X?”, screenshot the answer, and feel good about your brand sitting at number two. A week later you ask the same thing and a name you’ve never heard of is sitting where you were. Your brand is somewhere down the list now, or gone.

I’ve done this and felt the floor drop out. If the answer reshuffles every time I ask, what exactly am I supposed to be tracking? It’s tempting to throw up your hands and decide AI visibility (ie how the models mention, cite, and recommend you) just isn’t a measurable thing.

But hold on a second before you quit. “The list changed, so nothing here is measurable” is a leap — and I think it’s the wrong one. Something did change. The question is whether everything changed, or just the part you were staring at.

So let’s look at what actually moves when an AI answer wobbles, and what sits still underneath it. Three studies, and a much calmer conclusion than the screenshots suggest.

”It Named A Totally Different Brand Than Last Time”

Here’s the feeling, and it’s real: you can’t show your boss a number that’s different every Tuesday. If the ranking is a coin flip, a ranking dashboard is theater.

I get why people give up here. When the thing you measured this morning contradicts the thing you measured yesterday, the natural read is “this is noise, all of it.” Why pour budget into tracking a slot machine?

But that read smuggles in an assumption — that the only thing worth measuring was the rank. And rank is the single most fragile thing in the whole answer. Before you torch the idea of measuring, it’s worth asking how much of the chaos is the AI, and how much is just the one brittle number you picked to watch.

Reword The Question, And It Barely Moves

Start with the easy version of the worry: maybe the model is so jumpy that changing a single word in your prompt sends the answer flying.

A team led by Errica actually put a number on this (Errica et al., NAACL 2025). They measured prompt sensitivity — basically, how much a model’s predictions move when you reword the prompt without changing what you’re asking. Say the same thing two different ways; see if the model flinches.

Across several LLMs (large language models, ie the systems behind ChatGPT and friends), the paraphrase-induced swing came in at 3.2 to 10%. Not zero. Models aren’t rocks. But this is nowhere near “one reword and the whole thing collapses.”

So the surface phrasing isn’t the lever you feared it was. What moved results was the underlying meaning of the question, not the exact words. Which already tells you something useful: if you keep asking the same real question, you’re not standing on quicksand. The ground only shifts a few points when you shuffle the words around.

Maybe Your Measuring Stick Is Shaking The Result

Now the sharper finding, and the one that reframes the whole panic.

A study from Hua and colleagues asked whether all that apparent jumpiness is even the model’s fault, or something the evaluation method invents (Hua et al., EMNLP 2025). They ran seven LLMs across six benchmarks and twelve prompt templates — a lot of ways to ask, a lot of ways to score.

And much of the “variation” turned out to be an artifact of how they scored. When they used the old strict methods — log-likelihood scoring (ie grading by the probability the model assigns an answer) and exact-match (ie the answer only counts if it’s character-for-character identical) — the results looked all over the place. Switch to LLM-as-Judge (ie letting a capable model decide whether two answers actually mean the same thing), and the variance dropped while the rankings settled down.

Read that again, because it’s the crux. The evaluation protocol — the whole set of rules for how you score and compare the same target — was shaking the results harder than the model was. Same model, calmer ruler, steadier numbers.

So a big chunk of “AI recommendations are random” was never the AI. It was the measuring stick. And if your stick is bent, every reading off it is wrong in the same alarming direction.

In 2,961 Real Queries, One Number Held Still

Fine — but those are benchmarks. What happens when actual people ask actual AIs about actual brands?

Rand Fishkin’s team at SparkToro ran exactly that (SparkToro / Gumshoe, 2025). Six hundred people each asked three AIs the same kinds of brand questions — 2,961 queries in all — and the headline confirms your nightmare. Ask twice, and you get the identical list back less than 1% of the time. The order matches about 0.1% of the time. Rank is, for all practical purposes, a coin flip with a thousand sides.

But here’s the part that nobody screenshots. They also tracked how often a specific brand showed up at all — its appearance rate — and that number behaved. One high-visibility brand appeared in 97% of the 71 times its category came up. Not “ranked first 97% of the time.” Just there, reliably, almost every single ask.

So the order is noise and the presence is signal. “Where am I in the list” jumps around; “how often do I even appear” is trackable. That gap is the whole reason you shouldn’t quit — you were just reading the volatile column and ignoring the stable one sitting right next to it.

Conclusion: Stop Reading Rank, Start Counting Appearances

Here’s where I’ve landed. The single-shot answer was never a leaderboard, and treating it like one is what made you want to give up. One screenshot is one dice roll. You don’t judge a coin by one flip, and you shouldn’t judge your AI visibility off one answer.

What you can do is lock down the first two moves: how you ask, and how you tally. Pin the question (same real meaning, every time), then ask it over and over under identical conditions and count how often your brand appears. That appearance rate is the thing that holds still — the SparkToro numbers and the bent-ruler study both point at it. Rank wobbles; presence persists.

So when you’re picking a tracking tool, the thing I’d actually check is whether it tells you how it measures. Does it publish the repetition count, the appearance-rate method, and the comparison period it’s averaging over? A dashboard that just flashes a rank and won’t show its methodology is selling you the one number that was always going to shake — and maybe I’m wrong, but I wouldn’t let that be the thing steering real decisions.


Source

  • Errica, F., Siracusano, G., Sanvito, D., & Bifulco, R., “What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering”, NAACL 2025, arXiv:2406.12334
  • Hua, A., Tang, K., Gu, C., Gu, J., Wong, E., & Qin, Y., “Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs”, EMNLP 2025, arXiv:2509.01790
  • Fishkin, R. (SparkToro / Gumshoe), “NEW Research: AIs are highly inconsistent when recommending brands or products”, SparkToro, 2025, sparktoro.com
Share this article
Bluesky X
Back to all articles