“If The AI Is Taking It For Free, Let’s Shut The Door”
Somebody says a version of this in every content meeting eventually, usually about ten minutes in, usually with the whole room half-agreeing.
The awkward part is that neither side has a number. I didn’t either — when this landed on me last year the best I could manage was “well, what happens if we do?”, which is a question, not an answer.
But two pieces of measurement showed up this spring, and both of them went and looked at publishers who actually pulled the trigger.
One counted four million citations. The other tracked 30 newspapers before and after they blocked. Let’s take them in order.
Does Blocking AI Crawlers In robots.txt Stop You Getting Cited?
It mostly doesn’t. Among the top 50 news sites that were blocking a given bot, between 70.6% and 92.3% still turned up in AI citations anyway.
robots.txt is the note you leave on the front door saying which robots shouldn’t come in. It has no lock on it — it’s a request, and the bot decides what to do with the request.
The count comes from BuzzStream, which pulled about four million citations off 3,600 prompts across ten industries, published 8 April 2026, covering ChatGPT, Gemini, Google AI Overviews and Google AI Mode.
Sorted by which bot the site was blocking, the survival rate goes:
- ChatGPT-User, the one that fetches a page live during a chat: 70.6%
- OAI-SearchBot, OpenAI’s search index crawler: 82.4%
- GPTBot, OpenAI’s training crawler: 88.2%
- Google-Extended, the switch for Google’s training use: 92.3%
The pattern runs the wrong way round from the intuition. The harder a site pushes back on training, the more of its citations survive — and the only bot whose blocking makes a real dent is the one that goes and gets the page in real time.
Why Do Blocked Sites Keep Getting Cited?
Here’s the honest version: nobody knows yet, and the study says so itself.
The candidates it lists are the model already having learned the material, some other crawler that isn’t named in the robots.txt file, and the content coming back in through whoever republished it.
You can strike your entry from the alumni directory, but you can’t take yourself out of anyone’s memory. News in particular travels — the same story runs on a dozen syndication partners, so the model can describe your reporting perfectly well without ever touching your server.
Turn the same data around and look at it from the citation side, and the picture gets stranger. According to the same study, about 70% of everything ChatGPT cited came from sites blocking the fetch-type bots, and about 95% came from sites blocking the training crawlers.
The sites saying “don’t read us” are the mainstream of what gets read out loud.
I don’t think the mechanism matters much for a marketer, though, and it’s not settled anyway. What matters is the finding underneath it: closing the door doesn’t take your brand out of the answer. If your name is going to come up regardless, the useful question stops being whether to allow it and starts being what it says.
The Thing That Actually Fell Was The Visits
The second study is academic, and it measures the other side of the trade. Hangcheng Zhao at Rutgers Business School and Ron Berman at the Wharton School, University of Pennsylvania, tracked what blocking generative-AI crawlers did to a publisher’s traffic.
They compared weekly traffic at outlets that blocked against outlets that didn’t, over a main sample of 30 newspapers running from November 2022 to May 2024.
The method is difference-in-differences, which is a formal way of asking: the blockers changed by this much, the non-blockers changed by that much, so how big is the gap that only the blockers got?
Within six weeks of a block going up, the estimates land at −7.4% on SimilarWeb, −6.9% on Semrush and −6.5% on Comscore. Three independent traffic panels, all pointing at about a 7% drop.
Put the two studies side by side and you get the asymmetry that makes shutting the door such a bad deal: seven to nine citations in ten survive, and the visitors are the part that leaves.
How Much Of This Transfers To Your Own Site?
Less than you’d want, and I’d rather say that up front than let you carry these numbers into a meeting they don’t fit.
Both studies look at news publishers, whose content gets syndicated in a way a company’s own site never does. Your product pages don’t have twelve mirrors.
The BuzzStream side has real gaps too. It doesn’t disclose its measurement window, so there’s no guarantee the robots.txt state and the citation state line up in time; the citation data comes from a single tool (XOFU) with no third-party check on it; and the outfit that ran it sells digital PR software, which puts the conclusion “earn mentions elsewhere instead of blocking” comfortably close to its own business. That doesn’t make the count wrong, but it’s the kind of thing I’d want a reader to know about my own numbers.
The academic side is small — 30 newspapers in the main sample — and only the SimilarWeb estimate has a confidence interval that clears zero (ie, is statistically significant). By about 20 weeks out, even that one stops being significant. It’s also a preprint, meaning it’s been posted publicly before clearing journal review.
The authors flag two more limits themselves: the study window sits before AI-integrated search went mainstream, and consumption that happens entirely inside the AI’s answer never shows up in traffic panels at all.
And both studies run on mostly US outlets. Neither says anything about how you show up when someone asks in Japanese, or German, or Portuguese. That’s the part I’ve stopped guessing about and started measuring in my own market instead.
One more thing worth knowing if you go looking. An earlier version of this research got reported as a 23.1% monthly traffic decline, and that figure is still circulating. The same authors re-measured on weekly data and it came down to about 7% — so if you meet the big number, check you’re not reading last year’s draft.
Settle What The AI Says About You Before You Argue About The Door
The door turns out not to be the lever anyone in that meeting thinks it is. What the models say about your brand carries on with or without your permission, so the decision that actually moves anything is what that description contains.
The part you can touch today is small. Open your own robots.txt and read what it does with the AI bots — plenty of teams find a line nobody remembers adding, because a CMS or a starter template put it there.
Then start watching how you show up in the answers, separately from the blocking question entirely. Pick the three to five questions buyers in your category actually ask, put the same ones to the major assistants, and read what comes back. That’s enough to start with.
What happens after you let them in is the subject of the crawl-to-referral gap piece, and the file people add hoping to be read better is in the llms.txt one. This one sat one step earlier, at the point where you’re deciding whether to let them read you at all.
Shutting them out doesn’t erase the version of you the AI is describing. Maybe I’m wrong about how long that stays true — the crawlers keep changing. But if you can’t delete it, knowing what it says beats guessing.
Sources
- Hangcheng Zhao (Rutgers Business School), Ron Berman (The Wharton School, University of Pennsylvania), “Strategic Response of News Publishers to Generative AI”, arXiv:2512.24968, 2026-04-15, https://arxiv.org/abs/2512.24968