Your Website Is Their Training Data
Senior marketing leaders built their content strategy around Google’s referral traffic model. That model is breaking. Here’s what you do about it.
1. The Old Rules
For 25 years, the web ran on a handshake deal. Let us copy your content. We will send you traffic. You monetize that traffic. Everyone wins.
Googlebot was royalty. Ad bots were welcome. Analytics crawlers were family. Everything else got blocked at the door. Security teams kept the perimeter tight. Marketing teams optimized for the bots that helped. Both sides knew which bots were on their side.
It was not a perfect system. But it was a functioning one.
2. The New Rules
Then AI showed up. And we lost our minds.
Marketing leaders were told to optimize for LLMs. Structure your content for ClaudeBot. Make sure GPTBot can read you. AI will make you more discoverable. So we threw the doors open. We let in crawlers we never vetted, never measured, and never asked a single thing of.
We went from “block by default” to “please crawl my content, pretty please.” Nobody stopped to ask what we were getting in return.
3. The Problem
AI bots are gluttons. They crawl everything. Your blog posts. Your case studies. Your proprietary research. Your pricing pages. They ingest it all and serve it up inside their own interfaces. The user gets the answer without ever touching your site.
The numbers will make you wince. But the numbers are not the real story. The real story is why they work the way they do.
On June 3, 2026, Cloudflare CEO Matthew Prince posted data showing that automated requests had surpassed human ones for the first time in the internet’s history. Bots now generate 57.5% of all HTML web traffic. He had forecast this crossover for late 2027. It arrived 18 months early.
Bots just passed humans
In spring 2025, AI crawlers made up 22% of bot traffic. As of May 2026, that share had climbed to 52.3% for training alone, with another 35.7% mixed-purpose, and just 10.1% for live search indexing. The training slice is the critical number. When more than half of AI crawling exists to build model training corpora, the majority of that activity has no mechanism to send a visitor back to your site. Your content is consumed to improve a model that may later answer questions without ever linking to you.
What they're really doing
Only one slice of AI crawling has any mechanism to send a visitor back to you.
The extraction gap.
This is the single most important concept in the article. The crawl-to-referral ratio measures how many pages a crawler reads for every visitor it sends back. The spread across platforms is staggering.
According to Cloudflare Radar data from June 2026:
- Google: 5 pages per referral. The old bargain still works for them.
- DuckDuckGo: 2:1. Near parity.
- Microsoft (Copilot/Bing): 35:1. Still search-backed.
- Perplexity: 186:1. The most referral-efficient AI-first platform.
- OpenAI (GPTBot): 848:1. Reads nearly 850 pages for every visitor it sends back.
- Anthropic (ClaudeBot/Claude-User): 4,580:1. Reads 4,580 pages for every visitor it sends back.
The extraction gap
How many pages a crawler reads for every visitor it sends back. Pick a platform — the spread is staggering.
Let that sink in. Anthropic crawls your site nearly a thousand times more aggressively than Google does, relative to what it returns. Claude-User is now the second-busiest individual bot on the web at 11.3% of all verified bot traffic, trailing only GoogleBot at 14.2%. A single AI assistant agent out-crawls every search engine except Google.
The busiest bots on the web
A single AI assistant out-crawls every search engine except Google. AI crawlers now make up 20.3% of all verified bot traffic.
The mechanism is structural, not accidental. Googlebot crawls to index, then sends searchers to your page. The economics are symbiotic. Anthropic’s bot crawls to answer a question inside a chat interface, and the user rarely clicks through. The answer was already provided. The economics are extractive. The gap is not going to close. It is a feature of the product design, not a bug they will fix.
For every hour someone spends searching online, only 15 minutes is on the open web. The rest happens inside AI interfaces that never send the visitor back. HUMAN Security’s 2026 report found that agentic AI traffic grew roughly 7,851% year over year, nearly eight thousand percent, and automated traffic expanded approximately eight times faster than human activity.
But here is the part that matters most for a B2B marketing leader. According to PipeRocket’s 2026 study of 53 B2B SaaS brands, organic search still drives 91.3% of all traffic and generates 37 times more leads than every AI engine combined. AI referral traffic is real and growing fast, but organic is not dead. It is changing. And the brands that treat AI as a replacement for SEO instead of a parallel track are making a mistake.
This is not a security problem. Security teams can sleep easy. This is a marketing problem. And it is yours.
4. Two Bad Options
Option A: Keep the doors open. You stay visible in AI answers. Your brand gets name-dropped. But your content is raw fuel for someone else’s engine. You are feeding a machine that has no intention of paying the check. And that machine now accounts for over half of all automated requests on the open web.
Option B: Lock it down. You block AI crawlers. Your content stays yours. But you risk becoming invisible in an AI-first world where the majority of discovery is shifting inside chat interfaces. And once you block, do not expect AI companies to come knocking to negotiate.
Here is what the actual research says. And it is more nuanced than the headlines.
The paper, titled “Strategic Response of News Publishers to Generative AI” (Zhao, Rutgers & Berman, Wharton, April 2026), is the most rigorous study on this question. The authors constructed a publisher panel covering 30 major newspaper domains (CNN, NYT, WSJ, BBC, and others) and expanded certain analyses to the top 500 news-publisher domains by traffic. They used three independent traffic sources (SimilarWeb, Semrush, and Comscore’s household-level browsing panel) to cross-verify every finding. The empirical window runs from November 2022 through May 2024, stopping before Google introduced AI Overviews to avoid contaminating the causal analysis.
The headline finding is real but qualified. Among the top 30 publishers who blocked AI crawlers, total traffic declined by approximately 7% within six weeks of blocking. The effect is statistically significant across all three data sources. Critically, the Comscore panel, which tracks actual human browsing behavior, not server-side estimates, confirms the decline is in real human traffic, not just the mechanical removal of bot visits. The human-only traffic drop was 14%.
The authors propose two mechanisms. The more likely one is reduced brand exposure: when a publisher blocks a generative AI crawler, the LLM stops including that publisher as a source in query responses. Consumers who would have encountered the brand through an AI-generated answer instead encounter a competitor or no reference at all. Direct visits decline because brand recall weakens. The secondary channel is lost referral clicks from LLM citations, which the authors consider less significant given that AI referral traffic was still small before mid-2024.
But the size effect is the real story. When the researchers expanded their analysis to the top 500 publishers, the numbers flipped. For publishers ranked 51 to 100 by traffic, the effect weakened and lost statistical significance. For publishers ranked 101 to 500, the point estimate actually turned positive. Some mid-sized publishers experienced traffic increases after blocking AI crawlers.
Why would smaller publishers gain traffic by locking AI out? The most plausible explanation is the brand-awareness channel in reverse. Large publishers like CNN and the NYT have massive brand recognition. If they disappear from AI answers, a meaningful portion of their audience stops encountering them, and direct traffic falls. Mid-sized publishers do not have that brand pull. Their AI-answer exposure was primarily category-level: “what is X” queries where an AI would synthesize a comparison. When they blocked, they stopped losing those comparison queries to AI synthesis, and their existing organic audience was unaffected because those readers were searching for them by name anyway.
Some publishers reversed course within months of blocking, suggesting the traffic loss was material enough to override their concerns about AI training. Over 50 publisher-AI licensing agreements have been signed since 2023. The market is pricing this in real time.
Beyond access control, the study documented other defensive moves. Large publishers significantly increased rich content like video and interactive elements (68.1% increase) and ramped up advertising and targeting technologies (50.1% increase), creating experiences that AI tools cannot easily replicate.
Both options hurt. The question is which kind of hurt you can live with. And the answer depends entirely on who you are, who your buyers are, and how they find you.
5. A Third Option
All or nothing is a trap. The real play is tiered, intentional access based on what your content is worth to you.
a. Make yourself cite-worthy. Original research. Proprietary data. A framework the AI cannot safely make up. Structure your content so AI answers want to name you as the source. The goal is to be the footnote, not the training data.
b. Gate what matters. Let AI bots graze on your top-of-funnel thought leadership. Gate your pricing, your product specs, your competitive intel. You already do this for human visitors. Do it for bots too.
c. Build for the narrow funnel. AI answers can still drive traffic for specific types of queries. Unique data. Contrarian takes. Content an AI cannot plausibly generate on its own. Build for the narrow funnel, not the wide one. The PipeRocket data shows AI-referred visitors convert differently. They show up with less brand awareness (only 11.8% brand-name search intent versus 28.1% for organic), but they are further down the funnel. Separate research from mid-2026 found that AI search traffic converts at 14.2% compared to Google organic at 2.8%, a 5x advantage. The volume is lower but the quality is higher. Optimize for that.
d. Watch the licensing market. Over 50 publisher-AI agreements have been signed since 2023. Cloudflare is rolling out a Monetization Gateway for per-crawl payments. The infrastructure for getting paid is being built. You do not need to be the New York Times to participate. But you do need to be ready.
6. The Competitive Angle
Here is the question nobody wants to answer out loud. If you block ClaudeBot and your biggest competitor does not, who wins?
The answer depends on whether your buyers find you by name or by category.
A senior marketing leader researching “enterprise SEO platform” is searching by category. They do not have a brand in mind. If you are the only B2B brand visible in AI answers for that query, you own that discovery surface. If you blocked AI crawlers and your top three competitors did not, those three get every AI-generated comparison and recommendation. You are invisible. And the person who would have been your customer never knew you existed.
A CTO searching for “Atomic Glue case study” is searching by brand. They already know you. Blocking AI crawlers does not hurt that search. They will find your site through Google or direct navigation. The Rutger/Wharton data supports this: mid-sized publishers, which typically have stronger brand-direct traffic than category-discovery traffic, actually saw traffic increases after blocking AI crawlers. They were not losing category-level exposure because they did not have much to begin with, and they stopped AI from synthesizing their comparison content.
For B2B SaaS, the PipeRocket data is instructive. AI referrals make up under 9% of total traffic. But that under-9% pool converts differently. AI-referred visitors arrive with lower brand awareness (11.8% brand-name search intent versus 28.1% for organic), meaning they are discovery traffic, not brand-direct. They are the category searchers. If you block AI crawlers and your competitor stays open, that entire category-discovery pool goes to your competitor alone. For now, it is small. But AI search traffic converts at 14.2% compared to organic’s 2.8%. Every percentage point of AI referral growth compounds at a premium conversion rate.
A small pool at high quality
Organic still owns the volume. But AI-referred visitors convert at a premium — this is a parallel track, not a replacement.
The reverse is also true. If you stay open and your competitors lock down, you own the entire AI discovery surface for your category. That is a meaningful moat in a world where the second-busiest crawler on the web is now an AI assistant, not a search engine.
The zero-sum dynamic. AI visibility is a fixed pie in a way that search visibility was not. Google could send a dozen results to a query and everyone got traffic. AI answers typically cite one or two sources. If you are not in that answer, you do not exist in that query. The margin between being cited and being invisible is razor-thin, and every publisher who blocks increases the citation share for those who stay open.
The early blocking data illustrates the divergence. Major publishers like CNN and the NYT, brands so big that name searches dwarf category searches, lost 23% of total traffic when they blocked. Mid-sized B2B publishers with 50,000 to 500,000 monthly visits saw no effect or traffic gains.
Who loses when you block
The effect flips with size. Household-name brands bleed traffic; mid-sized publishers can gain it.
If you are a brand that people search for by name, blocking costs you more than it saves. If you are a brand that lives on category discovery, blocking costs you the entire AI channel.
Neither position is obviously right. But the strategic choice needs to be deliberate, not accidental. And it needs to be revisited quarterly, because the competitive landscape is shifting faster than any 12-month content plan can keep up with.
7. The Content Strategy Decision Framework
Here is the framework we use at Atomic Glue when a client asks what to do about AI bots. It applies to every piece of content you publish.
Content Type x Bot Tier x Action
The decision framework
All-or-nothing is a trap. Set access one content tier at a time.

Not sure where your content lands? Atomic Glue can pull your crawl-to-referral ratio from your own server logs in about an hour.
The crawl-to-referral audit. Pull your server logs for the last 90 days. Identify every AI crawler that hit your site. Compare crawl volume to referral traffic from each source. If a bot is crawling thousands of pages and sending zero referred visitors, that is a relationship you should reconsider.
The quarterly bot review. Add a standing quarterly agenda item. Review which bots are crawling, what they are consuming, and what traffic they are sending back. Adjust your robots.txt and your content strategy accordingly. This should take 30 minutes. If your agency cannot produce this data, ask why.
8. The Timeline
Here is what we see happening and when it matters for your decision.
The window is open
Now (mid-2026). Cloudflare has shifted the default to block AI crawlers for all new domains. Over 27% of websites are accidentally blocking AI crawlers without knowing it. Most B2B sites are running on default settings and have no idea what their bot traffic looks like. The blocking rate among top news publishers has climbed above 60%, versus under 10% among top retail domains. AI crawlers account for 20.3% of verified bot traffic and 57.5% of all HTML requests are now automated. The window for making a deliberate choice is closing faster than most marketing leaders realize.
6 months. Google’s mixed-use crawler problem comes to a head. Publishers are pushing for separation between search crawlers and AI training crawlers. If Google splits them, the decision framework gets cleaner. If not, the ambiguity persists. Meanwhile, the crawl-to-refer ratios on every major AI platform are worsening as training volumes grow faster than search volumes.
12 months. The licensing market matures. Pay-per-crawl models move from pilot to standard. The question shifts from “should I block AI bots?” to “what is the right price for my content?” Publishers who can produce proprietary, high-demand data will negotiate from strength. Publishers running default settings will accept whatever terms are offered.
18 months. Agentic Internet becomes the dominant paradigm. Over 57% of HTML traffic is non-human today. That number climbs. The brands that built their AI visibility strategy early will have a durable competitive advantage. The ones that waited will find their category keywords already claimed by incumbents who made the deliberate choice while the window was open.
9. Five Questions Every B2B Leader Should Ask Their Agency
You should not need to be the expert on this. Your agency should be. Here are the questions that separate the ones who understand this from the ones who are still running last year’s playbook.
1. What AI crawlers are hitting my site right now, what are they consuming, and how much traffic are they sending back?
If they cannot answer this with data from your actual server logs, they are flying blind. Not SimilarWeb estimates. Not Cloudflare dashboard screenshots. Your server logs. This is not a theoretical question. It is a reporting question.
2. What is my crawl-to-referral ratio by bot?
Not by source. By individual bot. Googlebot has a ratio of 5:1. GPTBot is 848:1. ClaudeBot is 4,580:1. Perplexity is 186:1. Treating them as one category means you are making policy decisions on bad information.
3. What content on my site is most exposed to AI substitution?
The answer is probably your comparison pages, your glossary content, and your “what is X” explainers. Those are the pages AI loves to synthesize. Your agency should know which pages those are before you ask. They should also know which pages drive your brand-name traffic versus your category-discovery traffic, because the blocking calculus is different for each.
4. If we block training crawlers but keep discovery crawlers, what changes in our visibility?
This is the nuanced question. Most AI companies separate these. Google does not. The Rutger/Wharton study found that the traffic loss from blocking was driven primarily by reduced brand exposure, not lost referral clicks. A good agency partner knows the difference and can model the impact of each approach on your specific content mix.
5. What is our competitive position on AI visibility?
If your top three competitors are visible in AI answers for your category keywords and you are not, that is a problem. If they have all blocked and you are the only brand showing up, that is an opportunity. Your agency should be monitoring this monthly, not annually.
10. What You Should Do Now
This is a marketing decision, not an IT decision. Own it.
Your content strategy needs a bot tier, just like it has audience tiers. Ask two questions about every piece of content:
- What do I need AI to know about me?
- What do I need humans to come here to see?
Start measuring your crawl-to-referral ratio. Not just your pageviews. Track which bots are eating what. If a bot is devouring your most expensive content and sending nothing back, that should be a deliberate strategic choice. Not a default setting you forgot to change.
Add the quarterly bot review to your calendar. It takes 30 minutes. The cost of not doing it is making a strategic decision by accident.
11. The Window Is Open
We are in a transition period. The agentic internet is here. The business models are forming in real time.
Google’s mixed-use crawler makes it hard to opt into search without also opting into AI. Cloudflare has declared a default block. Major publishers who blocked lost 23% of their traffic. Some mid-sized sites gained. The answer is not the same for everyone.
The brands that navigate this well will treat it as a relationship strategy, not a technical toggle. They will make deliberate choices about what to share, what to gate, and what to charge for. The ones that wait will find the choice has been made for them. And they will not like what they wake up to.
Atomic Glue helps companies build, optimize, and scale high-performing websites, applications, and digital marketing systems. We work with marketing teams, startups, and agencies across the United States. If you want to know what your crawl-to-referral ratio looks like, we can show you in about an hour.
