“We’re Well-Known In Our Industry, So The AI Should Get Us Right”
Ask the AI for your brand name and it answers. The mention rate on your visibility report sits above the industry average, too.
Mention rate is the share of AI answers that name your brand at all.
The usual move in a review meeting is to point at that bar chart and move to the next agenda item.
Checking whether you get mentioned first is the right order — no name in the answer means there’s nothing left to check.
But when you force the AI to attach a source URL to its answer and then actually open that URL, the page sometimes isn’t there. And apparently that happens more, not less, to well-known firms.
I read the paper that counted this citation by citation, and it changed how I read our own reports.
What Should You Actually Measure For A Well-Known Brand?
Whether the source URLs the AI cites actually open — on top of the mention rate.
The 50 well-known Hungarian firms are publicly listed companies and local subsidiaries of global brands, all internationally recognized. Across questions that required a URL or a document reference, 52.69% of their citations didn’t open. For the 50 domestic small and mid-sized firms, that figure was 37.87%. The gap: 14.82 points, in the opposite direction from what you’d expect.
That’s the reverse of the assumption at the top.
Zoltán Varga, at Neural Awareness in Budapest, checked every citation the AI produced for 100 B2B firms, one URL at a time (R).
It’s a preprint, posted 2026-06-19 and not yet peer-reviewed. The title page lists Varga as the sole author, at Neural Awareness and nowhere else.
The 100 firms were split into two tiers by how well-known they are.
- Well-known tier: 50 firms — publicly listed companies and local subsidiaries of global brands, all covered by international media
- Low-visibility tier: 50 domestic small and mid-sized firms with little international presence and little English-language coverage
Seven question types, all in Hungarian, covered things like policy, ISO certification, regulatory filings, and company registration — and every question required a URL or a document reference in the answer.
Two models answered: Claude Sonnet 4.6 and GPT-4o, neither with live web search — both answer from what they learned during training, not from a fresh lookup.
Each firm got the same 7 questions put to both models once, for 100 firms — 1,400 prompts total. Out of that, 2,062 citations came back, and every one was checked by a script, not judged by another AI. “Didn’t open” covers a URL that failed to return a normal page and a DOI (a registration number papers and documents get) that wasn’t found in the registry.
Low-Visibility Firms Got Fewer Citations In The First Place
Per question that required a URL, the share of answers that actually included at least one: Claude cited a source 85.7% of the time for well-known firms and 74.3% for low-visibility ones. GPT-4o cited a source 30.0% of the time for well-known firms and just 14.3% for low-visibility ones — meaning GPT-4o often declined to cite anything at all for the smaller firms.
Varga’s own reading: facing an unfamiliar firm, the model tends to back off rather than cite. That’s why the low-visibility tier’s failures skew toward “no citation” rather than “broken citation.”
So why did the well-known tier get more broken citations instead of fewer?
Varga names the pattern “the brand hallucination paradox.” According to him, the model knows a well-known firm well enough to construct a URL that looks right, without having learned enough to tie it to a page that actually exists.
That’s a reading, though, not a controlled test — this is an observation across two tiers, not an experiment where visibility was raised and the broken-citation rate was tracked as it moved. The tiers also differ in company size, and the paper doesn’t separate how much of the gap comes from visibility versus size.
Mixing In Regulatory Questions Surfaces The Failure You Actually Want To Catch
Broken-citation rate by question type: questions about the past three years of regulatory action or fines came out highest at 56.77%, with GDPR-policy questions close behind at 53.05%.
The company-registration question — used as a control — sat at 37.59%; the ISO-certification question (35.33%) didn’t clearly separate from that control.
Of the 946 citations that didn’t open, 478 returned a straightforward “page not found.” The domain itself existed; only the specific page didn’t.
That’s like being given the right train station but the wrong exit number.
Of those 478, 51.9% pointed not at the company’s own site but at Hungary’s competition authority, data-protection authority, central bank, or courts.
How many of your own prompt sets ask about regulatory history or GDPR compliance? If the answer is zero, you’re measuring with the exact question types this paper found least reliable left out.
This Number Is Hungarian, And The Questions Were Built To Force A URL
Requiring a URL in every answer is a design choice that pushes toward fabrication on purpose — all seven question types share that shape. Does that mean half of everyday citations, without that forcing, are made up too?
The paper doesn’t support taking it that far. The 52.69% figure comes with several conditions attached.
- Both models tested skip live search; measurement happened at one point in time, 2026-05-22
- Perplexity searches live and sits outside the main test. A small 2-firm side comparison put its broken-citation rate at 7.32% of 123 citations. The non-searching models ran 32–44% on that same side test
- 1,881 of the 2,062 citations came from Claude. GPT-4o’s tier-by-tier numbers aren’t reported separately, so that side of the split can’t be tracked from what’s given
- The rate is an upper bound. It counts everything that couldn’t be verified, and roughly 16.4% of citations marked “broken” were merely unreachable — blocked access, for instance — not confirmed as nonexistent
- The firms are Hungarian B2B companies, the questions are in Hungarian, and no Japanese-language measurement is in this paper
- It’s a preprint. It also counts citations from the same firm and question as separate data points, which the author himself flags as possibly making the “this gap isn’t chance” statistical test a bit more generous than it should be
Pasting 52.69% into your own forecast asks the number to travel further than its conditions allow. What this measurement supports is narrower: the two tiers failed in different ways, and that difference is worth checking for.
Decide Which Tier You’re In, Then Pick One Failure To Chase
The split ran on one line: publicly listed companies and local subsidiaries of global brands with international media coverage on one side, domestic firms with little international presence on the other. Firms in between weren’t part of the comparison.
If your brand clearly sits on the well-known side, I’d add source-checking to whatever you already track for mention rate — pull the citations an AI attaches to your name and actually open them, with regulatory and GDPR-style questions in the mix. Those were the question types that produced the most broken citations here.
If your brand sits clearly on the low-visibility side, the order flips: check whether your name and any citation show up at all before worrying about whether that citation opens.
The accuracy of what gets cited, once it does open, is the next layer — When AI Cites Your Brand, About Half The Time Something’s Off covers that. Switching which failure you’re chasing sometimes turns up a gap the mention-rate chart never showed.
Sources
- Varga, Z., “Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration,” arXiv:2606.21595, 2026-06-19, https://arxiv.org/abs/2606.21595