Back HexScope Lens

What Hundreds Of Real ChatGPT Questions Reveal About 'Best Tool' Prompts

POINT Key points
  • Real questions to ChatGPT run 71% longer than estimated prompts
  • Best-type prompts alone measure a conversation users rarely have

Is “Best CRM Software” Really How Anyone Asks?

If you’re testing your brand’s AI visibility, your prompt list probably looks something like “best CRM software” or “top project management tools”. Short, clean, category-shaped.

I’ve written lists exactly like that. They’re easy to standardize and easy to track over time.

But think about the last time you actually asked ChatGPT for a recommendation. Odds are you didn’t type two words — you typed a small paragraph, with your team size, your budget, and that one weird constraint your boss added last week.

If real questions look like that, then a test built entirely on “best X” prompts is measuring a conversation that mostly doesn’t happen. A few people have now measured that gap directly, so let’s look at what they found.

Real Questions Are 71% Longer (And Much More Personal)

Thomas Peham at Otterly AI compared hundreds of real ChatGPT prompts against the estimated prompts that tracking tools generate (the analysis).

The differences aren’t subtle:

  • Average length: 8.8 words for estimated prompts, 15.1 for real ones — 71% longer
  • Personal pronouns: in 18.8% of estimated prompts, 52.1% of real ones (2.8x)
  • Problem-oriented phrasing: 7.1% estimated, 21.1% real (3x)
  • First words: estimated prompts lean hard on “best”; real ones start with “what” and “I”

So real users aren’t asking for a ranking. They’re describing a situation — “here’s my team, here’s my budget, what actually works for us?” — which is less a search query and more a consultation.

It’s Not The Wording, It’s What’s Being Asked

Maybe you’re thinking: fine, but does phrasing really change the answer that much?

Here’s where a NAACL 2025 paper by Errica and colleagues helps (the study). They quantified how much LLM (large language model) accuracy shifts when you rephrase a question without changing its meaning: about 3.2 to 10%. Noticeable, but not the main event.

The main event is the meaning itself. Swapping “best” for “top” is cosmetic; swapping “rank these tools” for “which of these fits a 5-person team with no IT department” changes what kind of answer the model even tries to give.

And that second kind is the one real users keep asking.

Your Test Prompts Are Also A Thing To Test

There’s a third finding worth folding in. Mburu and colleagues studied AI-generated survey questions and reported they’re fast to produce but prone to verbosity and (the paper).

Which means the prompts in your tracking set are themselves an artifact worth checking. Ask an AI to “generate typical customer questions” and you’ll tend to get short, commercial, best-flavored ones — Otterly’s data shows exactly the direction of that drift.

So the fix isn’t “never generate prompts”. It’s: deliberately mix in the long, conditional, consultation-style questions the generator won’t hand you on its own.

Conclusion: Stop Testing Only “Best”, Start Testing The Consultation

Here’s where I’ve landed. If your AI visibility tracking runs entirely on short “best X” prompts, you’re measuring something real but narrow — and drifting away from the questions actual buyers ask, which come loaded with team sizes, budgets, and constraints.

Two things follow. First, build your prompt set in two layers: keep the short best-type prompts (they’re comparable over time), and add consultation-style ones with real conditions in them. Second, check your own content against those conditional questions — can anything on your site answer “which option fits this situation”, or only “here are our features”?

Maybe your best-type numbers are already fine. But the consultation is where the decision actually happens, and I think that’s the conversation worth watching.


Sources

  • Thomas Peham (Otterly AI), “Real vs Estimated Prompts: I Analyzed 100s of Real ChatGPT Queries”, 2026-02-03, Otterly AI
  • Federico Errica, Giuseppe Siracusano, Davide Sanvito, Roberto Bifulco, “What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering”, NAACL 2025, arXiv:2406.12334
  • Mburu, T. K., Rong, K., McColley, C. J., & Werth, A., “Methodological foundations for artificial intelligence-driven survey question generation”, Journal of Engineering Education, 114(3), e70012, 2025, DOI:10.1002/jee.70012
Share this article
Bluesky X
Back to all articles