Back HexScope Lens

ChatGPT And Claude Name Different Brands But Agree On Why You Lost

POINT Key points
  • Recommendations diverge two-thirds of the time, diagnoses agree 95%
  • Diagnose the stage you lose at before splitting your plan by provider

“So do we build one plan for ChatGPT and a second one for Claude?”

Ask a room to lay out AI-visibility work and this question arrives within ten minutes, every time.

I stalled on it for the better part of six months. Open any measurement dashboard and the brands ChatGPT names are simply not the brands Claude names. When the outputs disagree that badly, splitting the plan looks like the honest answer.

It also doubles the work. Which is why it’s worth checking whether the disagreement runs all the way down.

Somebody measured that — not the recommendations, but the reasons brands failed to get one.

Should You Build Separate AI Visibility Plans For ChatGPT And Claude?

Not before you diagnose. The brands the two providers recommend diverge about two-thirds of the time, but when both leave a brand out, the diagnosis of why matches 95.1% of the time (Will Jack and colleagues at Unusual.ai, posted May 2026, 215 commercial prompts, 7,763 joint failures).

So one diagnosis, run once, carries across both providers.

The diagnosis here means sorting your problem into a stage: the models can’t find you, or they find you and you don’t appeal, or they’ve filed you under a different job entirely. The research team sorted every joint failure into those three modes.

215 Commercial Prompts, 7,763 Cases Where Both Providers Left You Out

The team at Unusual.ai went after a practical fork: if your customers use both ChatGPT and Claude, do you optimise per provider (arXiv:2606.26116)?

The method is plain enough. Fire 215 commercial-context prompts at both providers, line up the recommendations, then pull the 7,763 “joint failures” — the runs where neither provider recommended the brand — and check whether the failure diagnosis agrees.

The divergence gets pinned down first. Cross-provider Jaccard similarity came in at 0.35, while re-running the same prompt against the same provider scores 0.50 to 0.61. The gap between providers is wider than the gap between two runs of the same provider.

The Less Famous You Are, The More The Diagnoses Agree

Agreement overall: 95.1%, confidence interval 94.3% to 95.7%. The breakdown underneath is where it gets useful.

  • Category-leader brands: 81%
  • Long-tail and regional brands: 99.6%

The more obscure the brand, the more both providers drop it for the same reason. Turn that around: once you’re well known, your failure modes start to diverge by provider.

Which tells you how long the shortcut lasts. While you’re still building recognition, one diagnosis covers both.

Different Answers, Same Diagnosis, Because The Routes Differ

The two providers reach a recommendation by visibly different roads. Anthropic answers from what it already knows 43-52% of the time; OpenAI does so 8-29% of the time.

One leans on memory, the other goes looking. Two different roads to the same question, so of course the names that come out don’t line up.

Yet the reason for the stumble is shared. Two people with completely different commutes can still both be late for the same reason: neither one left the house.

If you haven’t left the house, optimising two routes doesn’t help.

How Far This Number Travels

Four limits, and I’d say all four before this goes into a deck.

  • Two providers only, OpenAI and Anthropic. Gemini and Perplexity aren’t in it
  • US, UK and EU markets, 19 industries. No Japanese-language measurement
  • A single-day measurement, so drift over time is unmeasured
  • All four authors work at Unusual.ai, a vendor in AI brand visibility

That last one bears on the direction of the conclusion. “Your diagnosis is portable” is a comfortable finding for a company that sells diagnosis.

For what it’s worth, it doesn’t contradict the divergence we’ve written about before. What diverges is the output; the classification of causes is the shared part.

Diagnose The Stage Before You Split The Plan

The takeaway is about sequence.

Before you build a provider-by-provider matrix, insert one step. Work out whether you’re losing at discovery — the model never surfaces you — or at selection, where it surfaces you and picks a competitor.

That answer holds across OpenAI and Anthropic with roughly 95% confidence. And the less established your brand, the more reliably it holds.

Whether to split the plan is a question worth asking. Just ask it after the diagnosis, not before.


Sources

  • Will Jack, Noah Lehman, Keller Maloney, Sarah Xu (Unusual.ai), “Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation”, arXiv:2606.26116v1, posted 22 May 2026, arxiv.org (215 commercial prompts across OpenAI and Anthropic; 7,763 joint failures analysed. Cross-provider Jaccard similarity 0.35 against a same-provider re-run baseline of 0.50-0.61. Failure-mode diagnosis agreed 95.1% [94.3%, 95.7%], rising from 81% for category leaders to 99.6% for long-tail and regional brands. Anthropic recommended from training priors 43-52% of runs, OpenAI 8-29%. US/UK/EU markets, 19 industries, single-day measurement; all four authors are employed by Unusual.ai, a vendor of AI brand visibility measurement.)
Share this article
Bluesky X
Back to all articles