Back HexScope Lens

2,961 AI Answers Later, Rank Is Noise And Rate Is Signal

POINT Key points
  • Across 2,961 AI queries, the exact list matched under 1%, order near 0.1%
  • The fix is tracking appearance rate and standardizing how many times you measure

“Yesterday It Said A, Today It Says B” — And You Stop Trusting It

Ask an AI which brands it recommends, and you get one answer today and a different one tomorrow. Company A on Monday, Company B on Tuesday. Watch that happen a few times and the honest reaction is: if the ranking won’t sit still, how am I supposed to measure anything?

I think that reaction is exactly right. But a single ranking wobbling and the measurement itself being useless are two completely different things. Change the unit you look at, and a signal you can actually keep shows up.

So today let’s pull that “keepable” signal out of a genuinely large observation: 2,961 tries.

2,961 Queries In, The Single Rank Almost Never Reproduced

SparkToro and Gumshoe ran the test at real scale — 600 people, 2,961 queries in total — just to see how much the recommendations move. The result was blunt.

  • The exact same list reproduced less than 1% of the time.
  • The exact order matched about 0.1% of the time.

So reporting one answer screen as “this month’s ranking” is a shaky move. It’s like rolling a die once and deciding the die favors ones. The rank, on this evidence, is mostly noise.

But The Rate Held Steady, And The Gaps Stayed Comparable

Here’s the part you can’t skip, though. In the same data, the appearance rate held. A high-visibility brand turned up in 97% of 71 tries, for instance — so when you ask “how often does it come up” instead of “what place is it in,” the reproducibility is right there.

Funnier still, even prompts that weren’t semantically similar tended to return overlapping sets of brands. That lines up cleanly with Errica et al. (NAACL 2025), who found that meaning structure moves the output more than surface wording does. Change how you phrase it and the answer holds; change what you mean and it shifts.

The Real Contest Is Measurement Design, Not Model Comparison

The trap in practice is leading with “which model is smarter” and quietly skipping over the measurement conditions. Whether you’re running the same prompt 60 to 100 times or calling it after fewer than 10 changes how much you can actually read into the numbers.

One more thing matters here: reporting not just the average but the around it. Two brands can look different when their error bands are simply overlapping — that case is more common than it sounds.

Conclusion: Standardize The Conditions, Not The Rank Chase

If you’re tracking AI recommendations month over month, matching your measurement conditions beats agonizing over which way the rank ticked. What 2,961 repeated measurements really teach is this: because the AI’s answers are noisy by nature, you should read the appearance rate, not the single rank.

So when you pick a tool, check whether it publishes its repetition count and the logic behind how it aggregates the appearance rate. Same goes for your own internal reports — line up the question design, the model conditions, and the aggregation window so they match. Let that slip and you can’t tell a real improvement from a change in how you measured.

If you want the whole picture across seven related studies, this roundup on why appearance rate beats rank is a good front door to the rest.


Sources

  • Rand Fishkin (SparkToro/Gumshoe), “NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility”, 2026-01-27, sparktoro.com
  • Federico Errica, Giuseppe Siracusano, Davide Sanvito, Roberto Bifulco, “What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering”, NAACL 2025, arXiv:2406.12334
Share this article
Bluesky X
Back to all articles