Back HexScope Lens

ChatGPT And Perplexity Don't Read The Same Web. So Stop Optimizing Them The Same Way

POINT Key points
  • ChatGPT leans on Wikipedia, Perplexity leans on reddit
  • About 86% of AI citations come from sources you control

“Okay, it’s time we got serious about AI.” That line has started showing up in our meetings too — let’s tune the content so ChatGPT actually says our name.

But the moment you try to move, your hand stalls. “AI” isn’t just ChatGPT. There’s Perplexity, there’s Google’s AI Overviews, and you’ve got a finite amount of time and people. Where do you even start? It’s tempting to go looking for the one magic move that works on all of them at once.

I get the pull of that. When you’re short on hours, a single lever that fixes everything is exactly what you want to find.

But there’s a pile of data that throws cold water on the idea. The engines pull their sources from wildly different places. So let’s turn that into a priority list instead of a wish.

ChatGPT Is Reading Wikipedia, Perplexity Is Reading reddit

Start with a big count run by Semrush, the SEO tool company, over three months, looking at which sites get cited in AI answers, broken out by engine.

The lineups barely overlap. In ChatGPT’s citations, Wikipedia is dominant — for some models it grabs anywhere from 26% to 48% of the top-10 citation share, all on its own. One site, that much of the pie.

Peek into Perplexity and the lead actor changes completely. Here reddit — yes, the forum — is the standout, taking roughly 46.7% of the top-10 citations. Same word, “AI,” but reddit sits at about 11.3% in ChatGPT and around 21% in AI Overviews. The weighting is all over the place depending on the engine.

It’s a bit like ordering the same set menu at three diners and getting three different plates. Show up at every diner clutching one map labeled “AI strategy” and you might find it didn’t match the menu at any of them.

The Citations Pile Up On A Tiny Handful Of Domains

So the lineups differ by engine. Does that mean the citations are scattered across the whole web? Actually, the opposite.

5W, a research firm, pooled six large studies covering roughly 680 million citations and found the top 15 domains alone accounted for about 68% of AI citation share. The places these engines reach for are a much smaller world than you’d think.

The takeaway is pretty simple. Rather than blindly chasing more backlinks and more exposure everywhere, it’s faster to first pin down where the engine you care about is actually looking. Grind on reddit tactics while you’re trying to show up in ChatGPT and you’ll keep missing the target.

One thing worth flagging: these are all commercial counts from tool companies. Each defines a “citation” differently and covers a different window, so a number like 46.7% or 68% is best read as “roughly this lopsided,” not gospel. But the bigger pattern — the map differs by engine — keeps showing up across separate studies.

The Foundation Is Still The Pages You Own

Read this far and it starts to sound like the game is “how do I break into Wikipedia and reddit.” But step back one more notch and a different picture shows up.

Yext, which sells digital-presence management, analyzed 6.8 million citations across ChatGPT, Gemini, and Perplexity and found about 86% of the sources the AI pulled from were places the brand can control. The split: your own site at 44%, various listings at 42%.

“Listings” here basically means the places you register company info — a Google Business Profile, a MapQuest entry, that sort of thing. Together with your own site, both are places you get to shape the content of. Third-party forums and reviews, measured on a like-for-like basis, came out to just a few percent of the total.

And the engine gap shows up again. In the same Yext analysis, Gemini skewed toward the brand’s own site (52.1%), ChatGPT toward listings (48.7%), and Perplexity spread across a more varied mix of sources. Even at the level of what kind of source gets cited, each engine has its own habit.

Now, Yext sells exactly this kind of management tool, so “sources you can control matter most” happens to be a conclusion that helps their business. Read it with that discount in mind. But at 6.8 million citations with a disclosed method, it’s plenty useful as a directional signal.

Conclusion: Pick The Engine, Then Lock Down Your Own Ground First

So here’s what this run of data actually tells you to do.

The main thing: don’t lump “AI optimization” into one bucket. ChatGPT and Perplexity don’t cite the same cast of sources, so the smart move is to pick the engine you want to show up in and shape your work around what that engine reads. Chasing a single move that works on all of them wastes the little you’ve got. Read the maps one at a time.

On top of that, start with the ground under your feet. If most of the citations come from your own site and your listings — the places you can move — then get that info accurate and current first. That’s the foundation that tends to pay off no matter which engine you’re courting. The third-party stuff, Wikipedia or community, comes next, tuned to the engine you’re actually chasing.

And one more. If the citation sources swing this hard by engine, then not measuring “which AI cites us right now, and through what” leaves you blind to whether your moves are even landing. A rankings report won’t draw this map for you. Watch your per-engine citation picture on an ongoing basis, and only then can you actually set priorities. The entrance to AI work, I think, is closer to that plain question of “where do we stand right now” than to any clever tactic.


Sources

  • [R] Semrush (2026), “The Most-Cited Domains in AI: A 3-Month Study”, semrush.com
  • [R] 5W (2026), “AI Platform Citation Source Index 2026” (via PR Newswire), PR Newswire
  • [R] Yext, Inc. (2025-10-14), “Yext Research: 86% of AI Citations Come from Brand-Managed Sources”, yext.com
Share this article
Bluesky X
Back to all articles