The Report Cited Our Own Page For Something We’d Never Written
The report a deep research agent hands back looks like something a consultancy invoiced for: headings, a tidy summary, a numbered list of sources sitting underneath. Somewhere in that summary is a sentence about the company that isn’t anywhere on its site, and the link propping that sentence up is the site itself.
I opened the first one. The page loaded, the topic was right, and that was the last link I opened — checking by hand usually stops there, and I doubt I’m the only one who stops.
But six people did go through reports like that, one link at a time, and the check I skipped is the one that comes apart.
Do The Sources An AI Cites Actually Back Up What It Says About You?
Over 94% of those links open, and in over 80% of cases the page behind them sits on the same topic as the sentence doing the citing. Whether that page establishes the sentence as fact runs from 39% to 77% on the frontier models. Three checks on one citation. The third turns out to be a different animal from the first two.
The numbers come from a paper posted on 7 May 2026 by six people at PricewaterhouseCoopers, who pushed 130 research requests through 14 models and then reopened the citations the reports came back with, one at a time. Those requests were lifted from two public English-language benchmarks.
A deep research agent is the “go away and look this up properly” mode you’ve seen in ChatGPT and Gemini: the model reads dozens to hundreds of pages by itself, then writes a report with the links attached. And the reports in this study are that, and only that. According to the authors, the grading was done by a model working from a rubric, calibrated against 50 to 100 human reviews. Nobody sat down and read all of them, and I’d hold on to that once the percentages start sounding exact.
Your Visibility Report Is Counting The Outermost Layer
One citation, sitting under one sentence of a report. The paper puts three separate questions to it:
- Does the link open at all?
- Is the page on the other end about the same thing as the sentence?
- Does that page establish the sentence as fact?
When a visibility report puts a citation count next to your brand, that’s usually the first question, counted. Counting the second one as well shows that the model treated the page as the ground for what it wrote. That’s worth having, and it still says nothing about how the company was described.
Only the third question looks at that. I had been reading all three as one number. It’s the difference between checking that a phone number rings and checking who picks up: the page is real, the topic is right, and the number in the AI’s report still doesn’t match the number on your own page. Stack up the first count and the third stays exactly where it started, which is out of sight.
That last citation you checked — how far down the three did it get?
Turning The Searches Up To 150 Left The Link And Topic Checks Above 92%
The same team ran a second experiment on two models, winding the number of searches each one was allowed from 2 up to 150. Fact-checking fell away as the depth went up. Averaged across the two models it came down by about 42 percentage points. GPT-5.4 went from 79% to 17%; Claude Opus 4.6 from 80% to 58%.
On the two models they tried, the deeper the search went, the less of the report its own citations actually backed. You’d expect the opposite, wouldn’t you — I certainly did. And through all of that the first two checks held: at every depth, links opened and topics matched above 92%. One layer moved; the two outside it sat still.
What you can move from your side isn’t the model’s reading depth — it’s whether one sentence of yours survives being lifted out of the page on its own.
Since the paper didn’t test whether writing that way helps, take this next bit as mine rather than theirs. A sentence tends to hold up alone when it carries three things at once:
- the number itself (“up 32%”)
- when it’s from (“August 2026”)
- what it’s measured against (“year on year”)
“Shipments in August 2026 were up 32% year on year” can be quoted on its own and still say what it said. But “Up strongly on last year” needs the paragraph around it, and the paragraph is the part that doesn’t travel.
Where To Stop When You Carry These Numbers Home
What got measured is the citations inside research reports, written by agents in deep research mode. An ordinary chat answer sits outside it. And all 130 research requests came out of two public English-language benchmarks, so nothing here speaks to how the same thing goes when the question gets asked in another language.
PricewaterhouseCoopers is an accounting and consulting firm with a business helping companies adopt AI (which I kept in view the whole way through their paper about AI citations falling short). And it’s a preprint; no peer review yet.
The 14 models are specific versions, frozen where they stood when they were evaluated. The authors add a limit of their own. Web citations decay: pages come down and what’s on them changes, so a link that opened during the measurement may not open now, or may no longer say what it said then.
Strip all of that out and what’s actually left for your own brand?
So what I’d carry home is the frame rather than the figures: three layers, counted apart.
Open Ten Of This Month’s Citations And Write Down How Many Checked Out
Take the citations your brand picked up in AI answers this month and draw ten of them at random. Going by the ones that catch the eye gets a flattering sample, which is the whole reason for drawing them blind. Then open them. Read what each one says about the company, hold that against what your own page says, and note how many of the ten matched.
That figure isn’t the count of links that opened, and it isn’t the count of pages on the right topic. It’s the third layer, measured once, by hand — the only version of that number I’d trust for my own brand. Ten is a small sample, the way a spot check on one warehouse shelf is a small sample, and it still tells you whether the label matches the box.
If the count is already split into mentions and citations, this is the column that goes next to those two (we went through that split here).
A month where citations jump is a month somebody reads as things going well. But whether the accuracy came along with them is a separate question, and the tally column has never been able to answer it. Ten opened citations can, which is why I’d rather have those ten than another clean-looking total.
Sources
- Onweller, H. et al. (Commercial Technology and Innovation Office, PricewaterhouseCoopers, U.S.), “Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents”, arXiv:2605.06635, 7 May 2026 (limitations: a preprint, not yet peer-reviewed. The scope is citations inside research reports written by deep research agents, not ordinary chat answers. The 130 research requests come from two public English-language benchmarks. Topic-relevance and fact-check verdicts come from a rubric-based model judge calibrated against 50 to 100 human reviews, so not every citation was read by a person. PricewaterhouseCoopers has a business in AI adoption support. The 14 models are the versions current at evaluation time, and the authors note that web citations decay over time)