“The Citation Number Went Up And Nothing Else Did”
“Citations are up, though.”
That sentence lands in the monthly review, a finger goes to the chart, and the room half-nods. The line does go up.
But the pipeline looks the same as last month. So does branded search.
I used to answer this by widening the date range (which is roughly the move of replacing the batteries in a scale that says you haven’t lost weight).
The problem was the unit. If what you’re counting is too coarse, a longer window just gives you more of the same coarseness.
Then somebody cut that unit in half and counted again.
Being Cited, Being Used: Where’s The Difference?
Getting picked as a source and ending up in the substance of the answer are two separate stages. Across 602 controlled prompts and 21,143 citations, ChatGPT cited fewer sources than the others while the influence per retrieved page came out substantially higher (the paper). Perplexity and Google cited more sources on average.
The work is by Zhang Kai, He Xinyue and Yao Jingan, posted to arXiv on April 28, 2026.
Citation selection is the first stage: the model decides whether to search at all, then decides which pages to name as sources. That’s roughly where today’s AI visibility tools stop counting. Citation absorption is the second: how much of that page made it into the substance of the answer. It separates a page whose name sits in a source list from a page that supplied the skeleton of the reply.
They worked from a public dataset called geo-citation-lab, covering three families of systems: ChatGPT, Google’s AI Overview and Gemini, and Perplexity. From 21,143 valid citations and 18,151 pages they managed to retrieve, they pulled 72 features (length, structure, and so on) and compared.
What A Bar Chart Counted By Raw Citations Swallows
So here’s how it came apart. Breadth and depth separated cleanly.
- Perplexity and Google cite more sources on average (broad)
- ChatGPT cites fewer, but the influence per page is substantially higher (deep)
Which means one ChatGPT citation and one Perplexity citation carry different weight inside the answer — and that’s the part that actually costs you something. Line up your per-model visibility as bars counted by raw citations, and the weight difference vanishes entirely.
Picture the report: tall bar for Perplexity, short one for ChatGPT. Some of that gap is how highly each system rates you. Most of it is probably just Perplexity handing out more citations.
These days a per-model bar chart only starts meaning something to me once I’ve checked what the vertical axis counts.
What Did The Pages That Got Used Have?
The team sums up the shared traits of high-influence pages as four things:
- they’re long
- they’re structured
- they line up semantically with the intent of the prompt
- they’re loaded with extractable evidence: definitions, numeric facts, comparisons, procedures
Why those four? Probably because the model’s way of working looks less like “read, understand, summarize” and more like “pull out the usable fragments.” The more liftable pieces a page has lying around, the easier it is to build the skeleton of an answer out of it. At equal length, a run of definitions and numbers carries further than a run of opinions. (Turn it over and it gets uncomfortable: a page that reads beautifully, but is thin on cuttable evidence, can get cited and still stay outside the answer.)
The fourth one I was grateful for, because you check it by opening your own page. Is the term defined anywhere? Is there a concrete number? Is there a comparison with the alternatives? Do the steps run in order?
How Far Can You Lean On These Four?
For one, it’s a preprint — nobody outside has reviewed it — and the authors’ institutional affiliation isn’t stated, so on conflicts of interest there’s nothing to judge from. (It doesn’t look like vendor research, though the independence can’t be verified either.)
And the central result is descriptive statistics; the authors themselves call it a descriptive finding. In plain terms: they photographed what already-cited pages looked like, without changing a single one to see what happened. So what you can say stops at “the pages that got used were structurally thick.” The causal claim, structure your page and you’ll be absorbed, is never made.
The corpus is 602 controlled prompts, a bounded set, and none of it was measured in Japanese. The systems covered are ChatGPT, Google’s, and Perplexity; Claude and Copilot stay out.
So these four traits work as a lens for rereading your own pages, held with the caution a descriptive finding deserves.
Read Every Citation In Two Stages, Picked Then Used
Keep counting citations. How often you get picked is a genuine first-order signal that the models can find you. Add one column beside it, though, and the same number changes meaning: that column is the trace of the citation reaching the substance of the answer.
If you want to move something today, open a single important page. Of definition, number, comparison, and procedure, add whichever is thinnest. That’s enough for today; nobody’s asking you to assemble all four in an afternoon.
And drop the per-model comparison built on raw citation counts. ChatGPT and Perplexity hand out citations by different rules, so the moment you put them on one scale you’re reading them wrong. Split by model, and compare yourself against your own previous month.
From “how many times did our name come up” to “how far into the answer did we get.” Someone stopped treating one citation as simply one citation, which is a good sign for where this goes.