Back HexScope Lens

Is The AI Actually Inconsistent, Or Just Badly Measured?

POINT Key points
  • Individual AI answers wobble, but brand appearance rates stay stable
  • Much of the measured flakiness comes from the eval method itself

“But It Says Something Different Every Time”

Roll out AI at any company and someone will say it within a week: “I asked the same thing twice and got two different answers. How am I supposed to trust this?” It’s a completely reasonable reaction.

I’ve had the same experience — same question, different phrasing back, and a vague feeling that the ground was moving under me.

But here’s the trap: concluding “AI is unusable” from that experience alone. Recent research suggests that some of this wobble isn’t a flaw in the model at all. It’s inflated by the way we measure it. And once you can split those two apart, your options get a lot more concrete.

Prompt Sensitivity Isn’t Only Measuring The Model

Start with “prompt sensitivity” — how much the output moves when you reword the question. It’s usually treated as a model weakness, full stop.

Hua et al. went a step further. Instead of just observing the sensitivity, they asked whether the evaluation method itself was manufacturing some of it. They ran 7 major LLMs across 6 benchmarks and 12 prompt templates, trying to separate the model’s wobble from the ruler’s wobble.

That’s a question most benchmarks quietly skip.

A Lot Of The Wobble Is A Measurement Artifact

The finding: a substantial share of the observed sensitivity may be an artifact of the evaluation method — the ruler’s quirks making the wobble look bigger than it is.

Traditional evals lean on log-likelihood scoring (checking how probability mass is distributed) or strict answer matching (exact string match). Both are easy to automate, and both will happily mark two answers that mean the same thing as different (the “OK” vs “sure, that’s fine” problem).

So the authors also scored answers with an LLM-as-Judge, evaluating at the level of meaning. Result: performance variance shrank and model rankings got more consistent. The takeaway isn’t “AI is secretly stable” — it’s that what you count as a correct answer changes how much instability you see.

Brand Recommendations: Lists Wobble, Appearance Rates Don’t

Now for data much closer to marketing. SparkToro and Gumshoe had 600 volunteers run recommendation prompts against ChatGPT, Claude, and Google’s AI — 2,961 runs in total. The results are pleasantly counterintuitive:

  • The chance of getting an identical list twice was under 1%
  • The chance of the same list in the same order was about 0.1%
  • One high-visibility brand appeared in 97% of its 71 runs
  • Major brands held steady at 55–77% Visibility% (appearance rate)

So a single ranked list is nearly random. But the appearance rate across repetitions is remarkably steady.

One more detail I liked: the 142 human-crafted prompts in the study had a semantic similarity of just 0.081 — people asked in wildly different ways — and the answers still converged on similar sets of brands. Individual lists wobble; the aggregate tendency shows through.

One answer is noise. The appearance rate across many answers is signal. That’s the practical reading.

Conclusion: Track Appearance Rates, Not One-Off Rankings

Some AI wobble is real, and some of it is your measurement instrument making things look worse. Which means a single ranking snapshot is a dangerous KPI — you can’t tell luck from tendency.

For marketing work, the sturdier setup is to repeat same-intent prompts and treat Visibility% (the appearance rate) as your primary metric, with scoring that accepts semantically equivalent answers instead of demanding exact matches. Maybe I’m wrong about how far this generalizes — the research here is young. But it moves the conversation from “the AI is too flaky to use” to “the flakiness is a given; measure around it,” and that’s a much better conversation to be having.


Sources

  • Hua et al. (2025), “Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs”, arXiv
  • Rand Fishkin (SparkToro/Gumshoe) (2026), “NEW Research: AIs are highly inconsistent when recommending brands or products”, SparkToro
Share this article
Bluesky X
Back to all articles