“Can I Read This Like A Search Volume Or Not?”
Open almost any AI-visibility tool now and there it is: a count of how many people asked ChatGPT some particular thing last month. “Best project management tool” — 4,200 prompts. It looks exactly like the Keyword Planner screen you’ve been reading for a decade.
And I’ll admit my first reaction was “oh, handy.” If you’ve spent years ranking keywords by volume, a volume column for prompts is the most natural thing in the world; you already know what to do with it. You sort descending and start working.
But then you go looking for where the number comes from, and the floor gets soft. Nobody is standing behind every ChatGPT user with a clipboard. Somebody, somewhere, is estimating — and the estimate has a shape.
So let’s look at what that shape is, and how much of your judgment it can hold.
Where The Number Actually Comes From
Prompt search volume is a guess built from clickstream and panel data — a set of volunteers whose browsing gets recorded, then scaled up to stand in for everybody. That’s the whole trick, and everything else follows from it.
A critical review from jaeckert-odaniel.com picks the method apart, and it’s not gentle. Three holes, roughly:
- The panel is lopsided. It leans on desktop Chrome extensions, so mobile-app use mostly falls out — and panels skew tech-y, male, professional to begin with.
- A small sample gets stretched a long way. Thousands to tens of thousands of people stand in for a planet, and the error scales right along with the extrapolation.
- It borrows the search mindset wholesale. Asking a model to do something gets counted as “demand,” and Google Ads match-type thinking gets bolted onto sentences people write to a chatbot.
Think of it as reading the attendance list from one developer meetup and reporting the national average. You’d catch the problem instantly there. But dress the same move up in a familiar volume column and the — the part of real usage the method simply never sees — stops looking like a gap at all.
A Number Existing And A Number Being Right
Here’s the line from that review that stuck with me:
The reported figures look precise, but they carry a pseudo-precision that invites strategically significant errors.
“Pseudo-precision” is a number that displays more accuracy than the method behind it can back up. It’s the bathroom scale reading to two decimals while sitting three kilos off.
And the display does work on you. See “4,200 prompts a month” and your gut fills in a margin of error — plus or minus a hundred, maybe. That feels reasonable. It’s also invented. The real uncertainty here can run to an order of magnitude, which is a polite way of saying the second digit was never yours to read.
That’s the genuinely dangerous part. Not that the number’s wrong — plenty of useful numbers are wrong — but that it’s wrong while looking like the kind of thing you’d defend in a planning meeting.
Sometimes The Wobble Is In The Ruler, Not The Thing
Now for a study that rhymes with all this from a completely different direction. Hua and colleagues at EMNLP 2025 asked a sharp question about a different measurement panic (Hua et al., EMNLP 2025): when a model’s performance swings as you reword a prompt, is that the model wobbling, or the ruler? They checked seven major LLMs (large language models) across six benchmarks and twelve templates.
So here’s what they found:
- Most of the swing traced back to the scoring method — log-likelihood and exact-string matching.
- Switch to LLM-as-Judge (ie letting a capable model decide whether two answers mean the same thing) and the variance dropped while rankings got more consistent.
- The real difference between semantically equivalent prompts was smaller than anyone assumed.
Which means a lot of “LLMs are fragile” was never about the LLM. It was an artifact of how we graded it.
I think that lesson lands squarely on prompt volume too. “The number is unstable” and “the thing being counted is unstable” are separate claims, and they call for opposite responses. Blame the market for noise your instrument invented, and you’ll go fix the wrong thing with real money. That’s the failure I’d worry about most here — not being misled about demand, but being confidently misled about what to do next.
Conclusion: Demote It To A Direction, Then Cross-Check It
Prompt search volume isn’t worthless. It’s decent at direction — is interest in this topic rising or falling, are we in the same neighborhood as a competitor. Read it that way and it earns a small, honest place.
What it can’t hold is weight. Until a vendor tells you what the panel looks like and how much mobile it captures, that number has no business anchoring a budget split or sitting at the center of a scorecard. If they won’t answer the mobile question, that is the answer.
So the one habit worth building: never let it stand alone. Put it next to something independent — your appearance rate across models, your own inbound logs, Search Console — and only act where two of them agree. Maybe I’m being harder on it than it deserves; if a tool ever ties its panel to real logs and shows its work, I’ll happily upgrade it. But a number that can’t be cross-checked isn’t evidence. It’s just a very confident decimal point.
If you want to rethink the measurement design underneath all of this, ChatGPT Names A Different Brand Every Time. Measure Anyway. is the companion piece.
Sources
- jaeckert-odaniel.com, “Prompt search volume: Real data or all guessed?”, 2025-12-16, jaeckert-odaniel.com
- Andong Hua, Kenan Tang, Chenhe Gu, Jindong Gu, Eric Wong, Yao Qin, “Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs”, EMNLP 2025 Main Conference, arXiv:2509.01790
Related
- Does One Wrong Word Really Break ChatGPT’s Answer? — two studies on whether prompt phrasing really matters
- Your AI Rank Keeps Moving. ‘Appearance Rate’ Still Measures Your Brand. — why the model’s answer changes every time