More marketers are quietly typing their own company name into ChatGPT these days. “Who are the top vendors in my space?” — and then checking whether they show up. It’s the new mirror, and the urge to look into it is pretty understandable.
But anyone who’s tried it more than once has probably hit the same snag. Your brand was sitting at #3 a minute ago; you ask again, and the name’s just gone.
So the easy reaction is “AI is flaky, that’s just how it is.” And I’d have left it there too. But what if the flakiness follows a rule — and what if that rule is something you can actually control?
That changes the question. So let’s look at what happens when someone measures the wobble at scale: 14,000 queries’ worth.
How You’d Even Measure The Wobble
The team that ran this is the research group at Conductor, an AI-visibility shop. The scale is the part that makes it worth reading — 14,000 questions fired at the AIs in total, with the answers logged across four of them: ChatGPT, Perplexity, Claude, and Gemini.
The clever move was sorting the questions by intent. Intent is just the kind of thing the user is trying to do with the question (eg “which one should I buy” is purchase, “A or B?” is comparison, “what even is X?” is educational). They binned every question into seven of these intent types.
Then they measured the wobble two ways:
- Brand match rate: ask the same question twice, then take the share of brands that show up in both answers.
- #1 stability rate: the share of runs where the top pick — the AI’s first, headline recommendation — was the same both times.
Higher means steadier. So which kinds of questions were steady, and which fell apart?
Purchase Questions Are A Coin Flip; Comparisons Barely Move
The numbers came out in a clean staircase, which is more than I expected.
The shakiest by far was purchase intent. Its brand match rate was just 40%. In the team’s own framing: run a purchase prompt twice, and of the ten brands that turned up, only four showed up both times. The other six swapped in and out depending on the run.
The steadiest was comparison intent, with a #1 stability rate of 91%. Ask “A or B, which is better?” and you get the same top pick almost every time. Same AI, same question type — and yet purchase and comparison are living in two completely different worlds.
Educational intent had a weirder habit. The brand lineup itself was fairly consistent, but the #1 stability rate was only 30% — and 45-72% of these questions returned no brand name at all. Ask an AI “what is X?” and it would rather give you the general explanation than name a company, which honestly makes sense.
Why Buying Questions Are The Flakiest
Here’s the part that clicks. The wobble lines up with how crowded the question’s territory is.
A purchase question — “okay, but which one do I actually buy?” — opens onto a field of plausible candidates, lots of them, all roughly neck and neck. In a crowded, competitive space like that, the AI pulls a slightly different cast each time. Think of it like a grab bag: the winning brand’s in there some days, missing on others.
Comparison is the opposite. By the time someone asks “A vs B,” the field’s already been narrowed down to two. The contenders are fixed before the AI even answers, so there’s not much room for the result to drift. That’s why it holds steady.
There’s one more thing worth flagging: each AI has its own temperament. For the same question, ChatGPT averages around five brands per answer, while Gemini lists about 9.2. The more names an AI hands you, the more room there is for the lineup to churn — and Perplexity and Claude sat somewhere in between.
So which AI you measure on changes the baseline wobble before you’ve even chosen a question.
So What: Pin Down The Intent Before You Track
Pull it together and the takeaway is this. The AI’s answer doesn’t wobble at random — a big chunk of it is set by the intent of the question you asked. Purchase wobbles by nature; comparison barely moves. You can chase the same company name all day, and the picture you get still depends on how you asked.
So if you want to keep an eye on your own AI visibility over time, the first move is one thing.
Fix the intent of the question you measure with.
Tracking is just measuring under the same conditions, over and over. Ask purchase questions today and comparison ones next week, and you’ll never untangle whether a number moved because of something you did or because the question type just drifts. Want a stable signal to watch? Lock onto comparison. Want the messy, real-world churn? Use purchase. Either way, pick one on purpose and keep measuring it the same way.
And the second move matters just as much: don’t read too much into any single run. If your brand vanished on a purchase question one day, that might not be a failure — it might just be a naturally flaky corner. So go in expecting to read the average and the trend, not the snapshot, and you won’t get jerked around.
A small starting point: pick a handful of questions that describe you well, one or two per intent, and measure them on the same AI the same number of times, on a schedule. Then you’re reading your own numbers on your own yardstick — not somebody’s one-off “I asked the AI and it said…” post. That’s the first honest step toward living with AI visibility instead of being spooked by it.
Sources
- [R1] Jia-Rong Li (Conductor), “AI Brand Recommendation Study: Why Intent Type Predicts AI Output Consistency”, 2026. Conductor