“Okay, but what went up because of it?”
If you’ve ever presented an AI visibility report, you’ve heard some version of that question. You can say how often your name appeared. You can say it appeared more than last month.
But the follow-up — did anyone actually do anything — usually gets a shrug.
I assumed for a long time that this was a sequencing problem. Get the exposure number up first, and the business story shows up later.
Then a team tracked what people did after the AI answered, and the interesting split turned out not to be exposure at all.
Does Getting Mentioned By AI Actually Change What People Do?
Yes, though not evenly. Among people who hadn’t touched the brand at all in the previous week, an AI recommendation raised the odds of searching that brand by name within seven days by 4.3 percentage points, and the odds of visiting the brand’s own site by 2.4 points. Product pages on retail sites picked up 1.0 point.
That’s not a click-through count. It includes people who read the answer, went off, and typed the name into a search box themselves.
The study is a June 9, 2026 preprint (ie not peer-reviewed yet) from Michael Iannelli and Alan Ai at Scrunch AI, a company that sells AI visibility measurement.
The method: match conversation logs from ChatGPT, Claude, and Gemini against the browsing history of people who opted into a panel, at the individual level.
A clickstream panel is basically a record of the sites someone actually loaded — no survey, no “which brands do you recall seeing”, so none of the fuzz that comes with asking people to remember things.
The observation window runs seven days from the AI’s answer. The 95% confidence intervals (the range the estimate plausibly sits in) were 3.1–5.5 for branded search, 1.4–3.5 for own-site visits, and 0.3–1.7 for retail product pages.
One bit of arithmetic hygiene: these are percentage points, not percent. “3% became 7.3%” is a 4.3-point move, and reading it as “4.3% more” gets you a different (and wrong) number.
Recommended And Merely Listed Aren’t The Same Event
Here’s the part I’d underline. The authors sorted mentions into three stances — recommendation, neutral mention, and caution — and ran the same unexposed users through each.
A neutral mention moved branded search by 1.8 points, own-site visits by 1.1, and retail product pages by 0.3.
Against the recommendation numbers (+4.3 / +2.4 / +1.0), that’s roughly a third to a half of the movement for what looks, in a report, like the identical event.
Same brand, same week, same “one mention” in the tally — but whether the model vouched for you or just parked you in a list changed how many people went and did something by a factor of two or three.
If you’re counting mentions, those two land in the same column. It’s a bit like running a warehouse where tinned goods and fresh produce are both booked in by weight.
The Difference Holds Up Inside A Single Answer
The obvious objection is selection: maybe people who lean on AI are just more active shoppers generally, and the brands that show up are riding that.
So the authors compared, within the same AI response, the brands that got named against same-category brands that didn’t.
- Named in the answer: branded search +2.06 points, own-site visits +2.39 points
- Same category, not named: +0.20 points and +0.09 points
Same conversation, same context, same person. Whatever’s doing the work here, it isn’t the personality of AI users; it’s whether the name made it into that particular answer.
The two behaviors also travel together — branded search and own-site visits co-occurred 5.2 times more often than they would if they were independent. So they aren’t two separate wins. Track only one and you’re watching half of a single motion.
Pooling Mentions Hides Your Own Results
And this is where the measurement point gets expensive. Pool every mention together without sorting by stance, and the branded-search lift comes out at 2.08 points.
That’s less than half the 4.3 you get from recommendations alone. So a coarse tally does something worse than blur the picture. It quietly reports your own performance as smaller than it actually was.
There’s a second wrinkle worth knowing. Existing customers were already visiting the brand’s site 1.68 points above baseline before any mention, and the mention added only about 0.4 points on top.
The authors read that as AI recommendations landing mostly on people who weren’t already in the funnel, rather than pushing along a purchase someone had started. Which, for you, means these numbers sit in the new-customer conversation.
I think that’s the right read, but it’s the part of the paper I’d most want replicated.
What This Study Can And Can’t Tell You
Before anyone takes 4.3 to a planning meeting, the caveats, honestly:
- Scrunch AI sells AI visibility measurement, so “being recommended is worth money” is a conclusion with a commercial interest attached
- Panel size and per-cell counts are withheld as commercially sensitive, so nobody outside can check the effect sizes
- This is observational, not a randomized experiment (the authors lean on placebo windows at 14, 21, and 28 days prior, plus the within-answer control, to squeeze out confounds)
- The panel is opt-in and covers two English-speaking markets — not Japan, not most of the world. Three assistants, and Perplexity isn’t one of them
- It’s a preprint
Raw effects also look bigger for ChatGPT users than Gemini users — but that gap vanishes when you compare within the same person. So it’s a difference in who uses what, not in what the models can do to you.
Which is why I wouldn’t pin 4.3 up as a target. The transferable thing is the shape: recommendation and neutral mention are two to three times apart.
Add A “Were We Recommended?” Column
The count of mentions can’t stand in for the effect of them. One recommended appearance and one listed appearance are two very different events for the person reading, and averaging them together costs you more than half the measured lift.
So the fix is small and structural. Open the answers your key questions actually produce, and sort the ones that name you into recommended, neutral, and cautionary — then hold that split as a monthly rate on the same questions, because a single snapshot can’t tell you which way it’s moving.
That’s one column and one habit. If your current report has neither, I’d start there before adding another prompt to the list.
Sources
- Michael Iannelli, Alan Ai (Scrunch AI), “From Prompt to Purchase: How AI Brand Recommendations Move Consumers on the Open Web”, arXiv preprint 2606.10907, 9 June 2026, arxiv.org (conversation logs from ChatGPT, Claude and Gemini matched at the individual level against an opt-in clickstream panel, with behaviour observed over the seven days after the AI answer. Among users who had no contact with the brand in the previous week, a recommendation lifts branded search by +4.3 percentage points [95% interval 3.1–5.5], visits to the brand’s own site by +2.4 points [1.4–3.5] and retail product pages by +1.0 point [0.3–1.7]. A neutral mention gives, in the same order, +1.8 / +1.1 / +0.3 points. In the within-answer comparison, brands named in the answer record +2.06 points of branded search, +2.39 points of own-site visits and +0.51 points on retail product pages, against +0.20 / +0.09 / +0.16 points for same-category brands that were not named. Pooling every mention without distinguishing stance gives +2.08 / +2.37 / +0.52 points. Existing customers already sit +1.68 points above baseline on own-site visits before the mention, and the mention adds roughly +0.4 points. Branded search and own-site visits occur together 5.2 times more often than they would if they were independent. Raw effects are larger among ChatGPT users than Gemini users, but the gap disappears within the same person. Limits: the publisher, Scrunch AI, sells AI visibility measurement and improvement and so has a commercial interest in this conclusion. Panel size and per-cell counts are withheld as commercially sensitive, so the absolute size of the effect cannot be checked from outside. This is not a randomised experiment but an observational study; the authors address confounding with placebo windows at T−14, T−21 and T−28 days and with the within-answer control. The panel is opt-in and covers two English-speaking markets, so it does not measure any market outside them, and Perplexity is not among the assistants observed. It is a preprint that has not been through peer review. All effects are in percentage points, not relative changes)