Is The Tool Even Watching What People Actually Ask?
Sooner or later, anyone who buys an AI-visibility tool asks the same nervous question: is this thing tracking what real people type into ChatGPT, or a set of prompts it made up on my behalf?
It’s a fair thing to worry about. Most tools don’t sit behind millions of users watching them type; they estimate the prompts your buyers probably use, then run those. So if the guesses are off, you’re not measuring your market — you’re measuring the tool’s imagination of your market.
I think that instinct is dead right. But “the prompts are estimated” and “the numbers are useless” are two different claims, and the gap between them is where the whole thing gets interesting.
So let’s actually pull the two apart: where estimated prompts go wrong, and where they still earn their keep.
Real People Don’t Talk Like The Estimates
Start with how big the gap really is. Otterly AI ran the comparison, lining up hundreds of real ChatGPT prompts against the estimated ones. The differences aren’t subtle.
- Estimated prompts averaged 8.8 words; real ones averaged 15.1.
- First-person phrasing (“I”, “my”, “we”) showed up in 18.8% of estimated prompts and 52.1% of real ones.
- Problem-oriented wording (“I’m trying to…”, “we keep running into…”) hit 7.1% estimated versus 21.1% real.
Read that list and a picture falls out. Nobody types “best CRM.” They type “we’re a 5-person team on a tiny budget and our current CRM is a mess, what should we move to” — long, first-person, dragging their whole situation behind them. The estimate is short and shopping-list flat; the real thing is a confession.
That’s not a rounding error. It’s a different species of sentence, and any tool pretending its 8.8-word guesses are your customers is quietly lying to you.
The Rankings Barely Notice The Difference
Here’s the twist that saves the whole method. Despite those input styles being worlds apart, the brand rankings the two sets pulled back came out “quite similar.”
Which is strange, and worth sitting with. The words going in look nothing alike, but the names coming out mostly line up. So an estimated prompt isn’t a copy of reality — it’s more like a rough sketch of it, blurry on the details but right about who’s standing where.
And that’s the part you can use. If you only need to know whether you sit ahead of or behind a competitor, a sketch is plenty. What you can’t do is trust the fine print — the exact wording, the precise volume, the third-decimal-place stuff. Read the estimate for relative position, not for the literal question a customer supposedly asked.
The move is to stop asking “is this number accurate?” and start asking “is this gap between me and them consistent?” One of those is answerable. The other mostly isn’t.
”Prompt Search Volume” Is Dressed Up To Fool You
Which brings us to the metric that abuses all of this. A sharp critical review from jaeckert-odaniel.com takes apart “prompt search volume” — the feature that hands you a tidy count of how often people supposedly type a given prompt.
The trouble starts at the source. This data mostly leans on clickstream panels, which skew hard toward desktop Chrome and miss a big slice of mobile use. Then a biased panel gets extrapolated out to the whole market, and the error doesn’t cancel — it compounds.
So you end up with a number like “8,400 a month.” Looks precise. Isn’t. It’s the digital version of a bathroom scale that shows two decimal places (0.01 kg!) while being wildly miscalibrated — the extra digits feel like accuracy, but they’re just display. How many decimals a tool prints and how right it is are two unrelated things.
I’d keep this one on a short leash. Fine for “huh, maybe people do ask about this.” Dangerous the second it moves a budget line.
Conclusion: Check Where The Prompts Came From, Then Use Them As A Sketch
If there’s one habit to build, it’s this: before you trust any AI-visibility tool, find out where its estimated prompts actually come from. Real user logs? Scraped search terms? Something an AI dreamed up? Those aren’t interchangeable, and a tool that won’t tell you — or won’t say whether it captures mobile at all — hasn’t earned a headline number in your deck.
Then set your expectations to match what the data can actually do. Estimated prompts are a comparison aid, not a source of truth. Line them up next to your own first-party data (your analytics, your real inbound questions), read them for relative movement against competitors, and leave “prompt search volume” as a hint in the margin — never the KPI you present with a straight face.
The honest version isn’t “here’s exactly what people ask and how often.” It’s “here’s roughly where we sit versus the field, and whether that’s drifting.” Maybe I’m underselling the estimates a little — if a tool ever ties them to real logs, they get sharper. But until it shows you that, treat the guess as a guess, and you’ll be fine.
Sources
- Thomas Peham (Otterly AI), “Real vs Estimated Prompts: I Analyzed 100s of Real ChatGPT Queries”, otterly.ai
- jaeckert-odaniel.com, “Prompt search volume: Real data or all guessed?”, jaeckert-odaniel.com