[Atomic Glue](atomicglue.co)

Your Website Is Their Training Data

Senior marketing leaders built their content strategy around Google’s referral traffic model. That model is breaking. Here’s what you do about it.

1. The Old Rules

For 25 years, the web ran on a handshake deal. Let us copy your content. We will send you traffic. You monetize that traffic. Everyone wins.

Googlebot was royalty. Ad bots were welcome. Analytics crawlers were family. Everything else got blocked at the door. Security teams kept the perimeter tight. Marketing teams optimized for the bots that helped. Both sides knew which bots were on their side.

It was not a perfect system. But it was a functioning one.

2. The New Rules

Then AI showed up. And we lost our minds.

Marketing leaders were told to optimize for LLMs. Structure your content for ClaudeBot. Make sure GPTBot can read you. AI will make you more discoverable. So we threw the doors open. We let in crawlers we never vetted, never measured, and never asked a single thing of.

We went from “block by default” to “please crawl my content, pretty please.” Nobody stopped to ask what we were getting in return.

3. The Problem

AI bots are gluttons. They crawl everything. Your blog posts. Your case studies. Your proprietary research. Your pricing pages. They ingest it all and serve it up inside their own interfaces. The user gets the answer without ever touching your site.

The numbers will make you wince. But the numbers are not the real story. The real story is why they work the way they do.

On June 3, 2026, Cloudflare CEO Matthew Prince posted data showing that automated requests had surpassed human ones for the first time in the internet’s history. Bots now generate 57.5% of all HTML web traffic. He had forecast this crossover for late 2027. It arrived 18 months early.

Fig. 01 — The crossover

Bots just passed humans

57.5%of all HTML web traffic is now automated — the first time in the internet's history that bots out-number people.
Automated · 57.5%Human · 42.5%
Forecast: late 2027Arrived: mid-2026— 18 months early.
Source — Cloudflare (Matthew Prince), June 3, 2026

In spring 2025, AI crawlers made up 22% of bot traffic. As of May 2026, that share had climbed to 52.3% for training alone, with another 35.7% mixed-purpose, and just 10.1% for live search indexing. The training slice is the critical number. When more than half of AI crawling exists to build model training corpora, the majority of that activity has no mechanism to send a visitor back to your site. Your content is consumed to improve a model that may later answer questions without ever linking to you.

Fig. 02 — Why AI bots crawl

What they're really doing

Only one slice of AI crawling has any mechanism to send a visitor back to you.

88% has no path back to you
52.3%Training corporano path back
35.7%Mixed purposemostly no referral
10.1%Live search indexingcan send visitors
Source — Cloudflare, share of AI bot traffic, May 2026

The extraction gap.

This is the single most important concept in the article. The crawl-to-referral ratio measures how many pages a crawler reads for every visitor it sends back. The spread across platforms is staggering.

According to Cloudflare Radar data from June 2026:

  • Google: 5 pages per referral. The old bargain still works for them.
  • DuckDuckGo: 2:1. Near parity.
  • Microsoft (Copilot/Bing): 35:1. Still search-backed.
  • Perplexity: 186:1. The most referral-efficient AI-first platform.
  • OpenAI (GPTBot): 848:1. Reads nearly 850 pages for every visitor it sends back.
  • Anthropic (ClaudeBot/Claude-User): 4,580:1. Reads 4,580 pages for every visitor it sends back.
Fig. 03 — Crawl-to-referral ratio

The extraction gap

How many pages a crawler reads for every visitor it sends back. Pick a platform — the spread is staggering.

Anthropic — ClaudeBot / Claude-User
4,580pages crawled
per 1 visitor returned

Extractive — the busiest AI crawler on the web.

Relative to Google (5:1)
916×
Google’s crawl rate
Log scale · 1 → 10,000 pages per referral · click a row to compare
Source — Cloudflare Radar, June 2026

Let that sink in. Anthropic crawls your site nearly a thousand times more aggressively than Google does, relative to what it returns. Claude-User is now the second-busiest individual bot on the web at 11.3% of all verified bot traffic, trailing only GoogleBot at 14.2%. A single AI assistant agent out-crawls every search engine except Google.

Fig. 04 — Verified bot traffic

The busiest bots on the web

GooglebotSearch
14.2%
Claude-UserAI assistant
11.3%

A single AI assistant out-crawls every search engine except Google. AI crawlers now make up 20.3% of all verified bot traffic.

Source — Cloudflare Radar, share of verified bot traffic, June 2026

The mechanism is structural, not accidental. Googlebot crawls to index, then sends searchers to your page. The economics are symbiotic. Anthropic’s bot crawls to answer a question inside a chat interface, and the user rarely clicks through. The answer was already provided. The economics are extractive. The gap is not going to close. It is a feature of the product design, not a bug they will fix.

For every hour someone spends searching online, only 15 minutes is on the open web. The rest happens inside AI interfaces that never send the visitor back. HUMAN Security’s 2026 report found that agentic AI traffic grew roughly 7,851% year over year, nearly eight thousand percent, and automated traffic expanded approximately eight times faster than human activity.

But here is the part that matters most for a B2B marketing leader. According to PipeRocket’s 2026 study of 53 B2B SaaS brands, organic search still drives 91.3% of all traffic and generates 37 times more leads than every AI engine combined. AI referral traffic is real and growing fast, but organic is not dead. It is changing. And the brands that treat AI as a replacement for SEO instead of a parallel track are making a mistake.

This is not a security problem. Security teams can sleep easy. This is a marketing problem. And it is yours.

4. Two Bad Options

Option A: Keep the doors open. You stay visible in AI answers. Your brand gets name-dropped. But your content is raw fuel for someone else’s engine. You are feeding a machine that has no intention of paying the check. And that machine now accounts for over half of all automated requests on the open web.

Option B: Lock it down. You block AI crawlers. Your content stays yours. But you risk becoming invisible in an AI-first world where the majority of discovery is shifting inside chat interfaces. And once you block, do not expect AI companies to come knocking to negotiate.

Here is what the actual research says. And it is more nuanced than the headlines.

The paper, titled “Strategic Response of News Publishers to Generative AI” (Zhao, Rutgers & Berman, Wharton, April 2026), is the most rigorous study on this question. The authors constructed a publisher panel covering 30 major newspaper domains (CNN, NYT, WSJ, BBC, and others) and expanded certain analyses to the top 500 news-publisher domains by traffic. They used three independent traffic sources (SimilarWeb, Semrush, and Comscore’s household-level browsing panel) to cross-verify every finding. The empirical window runs from November 2022 through May 2024, stopping before Google introduced AI Overviews to avoid contaminating the causal analysis.

The headline finding is real but qualified. Among the top 30 publishers who blocked AI crawlers, total traffic declined by approximately 7% within six weeks of blocking. The effect is statistically significant across all three data sources. Critically, the Comscore panel, which tracks actual human browsing behavior, not server-side estimates, confirms the decline is in real human traffic, not just the mechanical removal of bot visits. The human-only traffic drop was 14%.

The authors propose two mechanisms. The more likely one is reduced brand exposure: when a publisher blocks a generative AI crawler, the LLM stops including that publisher as a source in query responses. Consumers who would have encountered the brand through an AI-generated answer instead encounter a competitor or no reference at all. Direct visits decline because brand recall weakens. The secondary channel is lost referral clicks from LLM citations, which the authors consider less significant given that AI referral traffic was still small before mid-2024.

But the size effect is the real story. When the researchers expanded their analysis to the top 500 publishers, the numbers flipped. For publishers ranked 51 to 100 by traffic, the effect weakened and lost statistical significance. For publishers ranked 101 to 500, the point estimate actually turned positive. Some mid-sized publishers experienced traffic increases after blocking AI crawlers.

Why would smaller publishers gain traffic by locking AI out? The most plausible explanation is the brand-awareness channel in reverse. Large publishers like CNN and the NYT have massive brand recognition. If they disappear from AI answers, a meaningful portion of their audience stops encountering them, and direct traffic falls. Mid-sized publishers do not have that brand pull. Their AI-answer exposure was primarily category-level: “what is X” queries where an AI would synthesize a comparison. When they blocked, they stopped losing those comparison queries to AI synthesis, and their existing organic audience was unaffected because those readers were searching for them by name anyway.

Some publishers reversed course within months of blocking, suggesting the traffic loss was material enough to override their concerns about AI training. Over 50 publisher-AI licensing agreements have been signed since 2023. The market is pricing this in real time.

Beyond access control, the study documented other defensive moves. Large publishers significantly increased rich content like video and interactive elements (68.1% increase) and ramped up advertising and targeting technologies (50.1% increase), creating experiences that AI tools cannot easily replicate.

Both options hurt. The question is which kind of hurt you can live with. And the answer depends entirely on who you are, who your buyers are, and how they find you.

5. A Third Option

All or nothing is a trap. The real play is tiered, intentional access based on what your content is worth to you.

a. Make yourself cite-worthy. Original research. Proprietary data. A framework the AI cannot safely make up. Structure your content so AI answers want to name you as the source. The goal is to be the footnote, not the training data.

b. Gate what matters. Let AI bots graze on your top-of-funnel thought leadership. Gate your pricing, your product specs, your competitive intel. You already do this for human visitors. Do it for bots too.

c. Build for the narrow funnel. AI answers can still drive traffic for specific types of queries. Unique data. Contrarian takes. Content an AI cannot plausibly generate on its own. Build for the narrow funnel, not the wide one. The PipeRocket data shows AI-referred visitors convert differently. They show up with less brand awareness (only 11.8% brand-name search intent versus 28.1% for organic), but they are further down the funnel. Separate research from mid-2026 found that AI search traffic converts at 14.2% compared to Google organic at 2.8%, a 5x advantage. The volume is lower but the quality is higher. Optimize for that.

d. Watch the licensing market. Over 50 publisher-AI agreements have been signed since 2023. Cloudflare is rolling out a Monetization Gateway for per-crawl payments. The infrastructure for getting paid is being built. You do not need to be the New York Times to participate. But you do need to be ready.

6. The Competitive Angle

Here is the question nobody wants to answer out loud. If you block ClaudeBot and your biggest competitor does not, who wins?

The answer depends on whether your buyers find you by name or by category.

A senior marketing leader researching “enterprise SEO platform” is searching by category. They do not have a brand in mind. If you are the only B2B brand visible in AI answers for that query, you own that discovery surface. If you blocked AI crawlers and your top three competitors did not, those three get every AI-generated comparison and recommendation. You are invisible. And the person who would have been your customer never knew you existed.

A CTO searching for “Atomic Glue case study” is searching by brand. They already know you. Blocking AI crawlers does not hurt that search. They will find your site through Google or direct navigation. The Rutger/Wharton data supports this: mid-sized publishers, which typically have stronger brand-direct traffic than category-discovery traffic, actually saw traffic increases after blocking AI crawlers. They were not losing category-level exposure because they did not have much to begin with, and they stopped AI from synthesizing their comparison content.

For B2B SaaS, the PipeRocket data is instructive. AI referrals make up under 9% of total traffic. But that under-9% pool converts differently. AI-referred visitors arrive with lower brand awareness (11.8% brand-name search intent versus 28.1% for organic), meaning they are discovery traffic, not brand-direct. They are the category searchers. If you block AI crawlers and your competitor stays open, that entire category-discovery pool goes to your competitor alone. For now, it is small. But AI search traffic converts at 14.2% compared to organic’s 2.8%. Every percentage point of AI referral growth compounds at a premium conversion rate.

Fig. 05 — Volume vs quality

A small pool at high quality

Organic still owns the volume. But AI-referred visitors convert at a premium — this is a parallel track, not a replacement.

Organic search
Share of traffic91.3%
Converts at2.8%
37× more leads than every AI engine combined.
AI referral
Share of traffic<9%
Converts at14.2%
5× the conversion rate — lower awareness, further down the funnel.
Source — PipeRocket 2026 B2B SaaS study (53 brands); AI conversion research, mid-2026

The reverse is also true. If you stay open and your competitors lock down, you own the entire AI discovery surface for your category. That is a meaningful moat in a world where the second-busiest crawler on the web is now an AI assistant, not a search engine.

The zero-sum dynamic. AI visibility is a fixed pie in a way that search visibility was not. Google could send a dozen results to a query and everyone got traffic. AI answers typically cite one or two sources. If you are not in that answer, you do not exist in that query. The margin between being cited and being invisible is razor-thin, and every publisher who blocks increases the citation share for those who stay open.

The early blocking data illustrates the divergence. Major publishers like CNN and the NYT, brands so big that name searches dwarf category searches, lost 23% of total traffic when they blocked. Mid-sized B2B publishers with 50,000 to 500,000 monthly visits saw no effect or traffic gains.

Fig. 06 — Traffic change after blocking AI crawlers

Who loses when you block

The effect flips with size. Household-name brands bleed traffic; mid-sized publishers can gain it.

← Traffic lostTraffic gained →
CNN, NYT & mega-brands−23%
Top 30 publishers−7% total · −14% human
Publishers ranked 51–100No significant effect
Publishers ranked 101–500Traffic gains
Source — Zhao, Rutgers & Berman (Wharton), "Strategic Response of News Publishers to Generative AI," April 2026

If you are a brand that people search for by name, blocking costs you more than it saves. If you are a brand that lives on category discovery, blocking costs you the entire AI channel.

Neither position is obviously right. But the strategic choice needs to be deliberate, not accidental. And it needs to be revisited quarterly, because the competitive landscape is shifting faster than any 12-month content plan can keep up with.

7. The Content Strategy Decision Framework

Here is the framework we use at Atomic Glue when a client asks what to do about AI bots. It applies to every piece of content you publish.

Content Type x Bot Tier x Action

Fig. 07 — Content type × bot access

The decision framework

All-or-nothing is a trap. Set access one content tier at a time.

Content typeLet AI crawl?Why
Foundational thought leadershipYesEstablish category authority. Low substitution risk.
Original research / proprietary dataConditionalLet AI cite you — gate the full dataset behind a form.
Product pages / pricingSelectiveBlock training crawlers. Allow discovery crawlers for search visibility.
Case studies with named clientsNoHigh substitution risk. Client relationships are proprietary.
Comparison / competitive contentNoAI synthesizes the comparison without sending anyone to you.
Atomic Glue moose mascot, thinking

Not sure where your content lands? Atomic Glue can pull your crawl-to-referral ratio from your own server logs in about an hour.

Source — Atomic Glue content strategy framework

The crawl-to-referral audit. Pull your server logs for the last 90 days. Identify every AI crawler that hit your site. Compare crawl volume to referral traffic from each source. If a bot is crawling thousands of pages and sending zero referred visitors, that is a relationship you should reconsider.

The quarterly bot review. Add a standing quarterly agenda item. Review which bots are crawling, what they are consuming, and what traffic they are sending back. Adjust your robots.txt and your content strategy accordingly. This should take 30 minutes. If your agency cannot produce this data, ask why.

8. The Timeline

Here is what we see happening and when it matters for your decision.

Fig. 08 — What happens when

The window is open

01Now · mid-2026The default flipsCloudflare blocks AI crawlers for new domains. 57.5% of HTML traffic is automated; 27% of sites block AI by accident.
02+6 monthsThe mixed-use fightPublishers push Google to split search crawlers from AI-training crawlers. Crawl-to-referral ratios keep worsening.
03+12 monthsLicensing maturesPay-per-crawl moves from pilot to standard. The question shifts from "should I block?" to "what’s my price?"
04+18 monthsAgentic internetNon-human traffic keeps climbing. Early movers hold a durable moat; latecomers find their keywords claimed.
Source — Cloudflare; HUMAN Security 2026 report; Atomic Glue analysis

Now (mid-2026). Cloudflare has shifted the default to block AI crawlers for all new domains. Over 27% of websites are accidentally blocking AI crawlers without knowing it. Most B2B sites are running on default settings and have no idea what their bot traffic looks like. The blocking rate among top news publishers has climbed above 60%, versus under 10% among top retail domains. AI crawlers account for 20.3% of verified bot traffic and 57.5% of all HTML requests are now automated. The window for making a deliberate choice is closing faster than most marketing leaders realize.

6 months. Google’s mixed-use crawler problem comes to a head. Publishers are pushing for separation between search crawlers and AI training crawlers. If Google splits them, the decision framework gets cleaner. If not, the ambiguity persists. Meanwhile, the crawl-to-refer ratios on every major AI platform are worsening as training volumes grow faster than search volumes.

12 months. The licensing market matures. Pay-per-crawl models move from pilot to standard. The question shifts from “should I block AI bots?” to “what is the right price for my content?” Publishers who can produce proprietary, high-demand data will negotiate from strength. Publishers running default settings will accept whatever terms are offered.

18 months. Agentic Internet becomes the dominant paradigm. Over 57% of HTML traffic is non-human today. That number climbs. The brands that built their AI visibility strategy early will have a durable competitive advantage. The ones that waited will find their category keywords already claimed by incumbents who made the deliberate choice while the window was open.

9. Five Questions Every B2B Leader Should Ask Their Agency

You should not need to be the expert on this. Your agency should be. Here are the questions that separate the ones who understand this from the ones who are still running last year’s playbook.

1. What AI crawlers are hitting my site right now, what are they consuming, and how much traffic are they sending back?

If they cannot answer this with data from your actual server logs, they are flying blind. Not SimilarWeb estimates. Not Cloudflare dashboard screenshots. Your server logs. This is not a theoretical question. It is a reporting question.

2. What is my crawl-to-referral ratio by bot?

Not by source. By individual bot. Googlebot has a ratio of 5:1. GPTBot is 848:1. ClaudeBot is 4,580:1. Perplexity is 186:1. Treating them as one category means you are making policy decisions on bad information.

3. What content on my site is most exposed to AI substitution?

The answer is probably your comparison pages, your glossary content, and your “what is X” explainers. Those are the pages AI loves to synthesize. Your agency should know which pages those are before you ask. They should also know which pages drive your brand-name traffic versus your category-discovery traffic, because the blocking calculus is different for each.

4. If we block training crawlers but keep discovery crawlers, what changes in our visibility?

This is the nuanced question. Most AI companies separate these. Google does not. The Rutger/Wharton study found that the traffic loss from blocking was driven primarily by reduced brand exposure, not lost referral clicks. A good agency partner knows the difference and can model the impact of each approach on your specific content mix.

5. What is our competitive position on AI visibility?

If your top three competitors are visible in AI answers for your category keywords and you are not, that is a problem. If they have all blocked and you are the only brand showing up, that is an opportunity. Your agency should be monitoring this monthly, not annually.

10. What You Should Do Now

This is a marketing decision, not an IT decision. Own it.

Your content strategy needs a bot tier, just like it has audience tiers. Ask two questions about every piece of content:

  • What do I need AI to know about me?
  • What do I need humans to come here to see?

Start measuring your crawl-to-referral ratio. Not just your pageviews. Track which bots are eating what. If a bot is devouring your most expensive content and sending nothing back, that should be a deliberate strategic choice. Not a default setting you forgot to change.

Add the quarterly bot review to your calendar. It takes 30 minutes. The cost of not doing it is making a strategic decision by accident.

11. The Window Is Open

We are in a transition period. The agentic internet is here. The business models are forming in real time.

Google’s mixed-use crawler makes it hard to opt into search without also opting into AI. Cloudflare has declared a default block. Major publishers who blocked lost 23% of their traffic. Some mid-sized sites gained. The answer is not the same for everyone.

The brands that navigate this well will treat it as a relationship strategy, not a technical toggle. They will make deliberate choices about what to share, what to gate, and what to charge for. The ones that wait will find the choice has been made for them. And they will not like what they wake up to.

Atomic Glue helps companies build, optimize, and scale high-performing websites, applications, and digital marketing systems. We work with marketing teams, startups, and agencies across the United States. If you want to know what your crawl-to-referral ratio looks like, we can show you in about an hour.

Jeff Walden
Jeff Walden, Managing Director

Jeff Walden is the Managing Director of Atomic Glue, where he works hands-on with clients on web development, SEO, and digital growth strategy.

Power upyourdigital world

Want to know what your crawl-to-referral ratio looks like? We can pull it from your own server logs in about an hour.

Atomic Glue moose mascot
# Your Website Is Their Training Data

Senior marketing leaders built their content strategy around Google's referral traffic model. That model is breaking. Here's what you do about it.

Author: Jeff Walden, Managing Director

**Senior marketing leaders built their content strategy around Google's referral traffic model. That model is breaking. Here's what you do about it.**

## 1. The Old Rules

For 25 years, the web ran on a handshake deal. Let us copy your content. We will send you traffic. You monetize that traffic. Everyone wins.

Googlebot was royalty. Ad bots were welcome. Analytics crawlers were family. Everything else got blocked at the door. Security teams kept the perimeter tight. Marketing teams optimized for the bots that helped. Both sides knew which bots were on their side.

It was not a perfect system. But it was a functioning one.

## 2. The New Rules

Then AI showed up. And we lost our minds.

Marketing leaders were told to optimize for LLMs. Structure your content for ClaudeBot. Make sure GPTBot can read you. AI will make you more discoverable. So we threw the doors open. We let in crawlers we never vetted, never measured, and never asked a single thing of.

We went from "block by default" to "please crawl my content, pretty please." Nobody stopped to ask what we were getting in return.

## 3. The Problem

AI bots are gluttons. They crawl everything. Your blog posts. Your case studies. Your proprietary research. Your pricing pages. They ingest it all and serve it up inside their own interfaces. The user gets the answer without ever touching your site.

The numbers will make you wince. But the numbers are not the real story. The real story is why they work the way they do.

On June 3, 2026, Cloudflare CEO Matthew Prince posted data showing that automated requests had surpassed human ones for the first time in the internet's history. Bots now generate 57.5% of all HTML web traffic. He had forecast this crossover for late 2027. It arrived 18 months early.

In spring 2025, AI crawlers made up 22% of bot traffic. As of May 2026, that share had climbed to 52.3% for training alone, with another 35.7% mixed-purpose, and just 10.1% for live search indexing. The training slice is the critical number. When more than half of AI crawling exists to build model training corpora, the majority of that activity has no mechanism to send a visitor back to your site. Your content is consumed to improve a model that may later answer questions without ever linking to you.

**The extraction gap.**

This is the single most important concept in the article. The crawl-to-referral ratio measures how many pages a crawler reads for every visitor it sends back. The spread across platforms is staggering.

According to Cloudflare Radar data from June 2026:

Let that sink in. Anthropic crawls your site nearly a thousand times more aggressively than Google does, relative to what it returns. Claude-User is now the second-busiest individual bot on the web at 11.3% of all verified bot traffic, trailing only GoogleBot at 14.2%. A single AI assistant agent out-crawls every search engine except Google.

The mechanism is structural, not accidental. Googlebot crawls to index, then sends searchers to your page. The economics are symbiotic. Anthropic's bot crawls to answer a question inside a chat interface, and the user rarely clicks through. The answer was already provided. The economics are extractive. The gap is not going to close. It is a feature of the product design, not a bug they will fix.

For every hour someone spends searching online, only 15 minutes is on the open web. The rest happens inside AI interfaces that never send the visitor back. HUMAN Security's 2026 report found that agentic AI traffic grew roughly 7,851% year over year, nearly eight thousand percent, and automated traffic expanded approximately eight times faster than human activity.

But here is the part that matters most for a B2B marketing leader. According to PipeRocket's 2026 study of 53 B2B SaaS brands, organic search still drives 91.3% of all traffic and generates 37 times more leads than every AI engine combined. AI referral traffic is real and growing fast, but organic is not dead. It is changing. And the brands that treat AI as a replacement for SEO instead of a parallel track are making a mistake.

This is not a security problem. Security teams can sleep easy. This is a marketing problem. And it is yours.

## 4. Two Bad Options

**Option A: Keep the doors open.** You stay visible in AI answers. Your brand gets name-dropped. But your content is raw fuel for someone else's engine. You are feeding a machine that has no intention of paying the check. And that machine now accounts for over half of all automated requests on the open web.

**Option B: Lock it down.** You block AI crawlers. Your content stays yours. But you risk becoming invisible in an AI-first world where the majority of discovery is shifting inside chat interfaces. And once you block, do not expect AI companies to come knocking to negotiate.

Here is what the actual research says. And it is more nuanced than the headlines.

The paper, titled "Strategic Response of News Publishers to Generative AI" (Zhao, Rutgers & Berman, Wharton, April 2026), is the most rigorous study on this question. The authors constructed a publisher panel covering 30 major newspaper domains (CNN, NYT, WSJ, BBC, and others) and expanded certain analyses to the top 500 news-publisher domains by traffic. They used three independent traffic sources (SimilarWeb, Semrush, and Comscore's household-level browsing panel) to cross-verify every finding. The empirical window runs from November 2022 through May 2024, stopping before Google introduced AI Overviews to avoid contaminating the causal analysis.

**The headline finding is real but qualified.** Among the top 30 publishers who blocked AI crawlers, total traffic declined by approximately 7% within six weeks of blocking. The effect is statistically significant across all three data sources. Critically, the Comscore panel, which tracks actual human browsing behavior, not server-side estimates, confirms the decline is in real human traffic, not just the mechanical removal of bot visits. The human-only traffic drop was 14%.

The authors propose two mechanisms. The more likely one is reduced brand exposure: when a publisher blocks a generative AI crawler, the LLM stops including that publisher as a source in query responses. Consumers who would have encountered the brand through an AI-generated answer instead encounter a competitor or no reference at all. Direct visits decline because brand recall weakens. The secondary channel is lost referral clicks from LLM citations, which the authors consider less significant given that AI referral traffic was still small before mid-2024.

**But the size effect is the real story.** When the researchers expanded their analysis to the top 500 publishers, the numbers flipped. For publishers ranked 51 to 100 by traffic, the effect weakened and lost statistical significance. For publishers ranked 101 to 500, the point estimate actually turned positive. Some mid-sized publishers experienced traffic increases after blocking AI crawlers.

Why would smaller publishers gain traffic by locking AI out? The most plausible explanation is the brand-awareness channel in reverse. Large publishers like CNN and the NYT have massive brand recognition. If they disappear from AI answers, a meaningful portion of their audience stops encountering them, and direct traffic falls. Mid-sized publishers do not have that brand pull. Their AI-answer exposure was primarily category-level: "what is X" queries where an AI would synthesize a comparison. When they blocked, they stopped losing those comparison queries to AI synthesis, and their existing organic audience was unaffected because those readers were searching for them by name anyway.

Some publishers reversed course within months of blocking, suggesting the traffic loss was material enough to override their concerns about AI training. Over 50 publisher-AI licensing agreements have been signed since 2023. The market is pricing this in real time.

Beyond access control, the study documented other defensive moves. Large publishers significantly increased rich content like video and interactive elements (68.1% increase) and ramped up advertising and targeting technologies (50.1% increase), creating experiences that AI tools cannot easily replicate.

Both options hurt. The question is which kind of hurt you can live with. And the answer depends entirely on who you are, who your buyers are, and how they find you.

## 5. A Third Option

All or nothing is a trap. The real play is tiered, intentional access based on what your content is worth to you.

**a. Make yourself cite-worthy.** Original research. Proprietary data. A framework the AI cannot safely make up. Structure your content so AI answers want to name you as the source. The goal is to be the footnote, not the training data.

**b. Gate what matters.** Let AI bots graze on your top-of-funnel thought leadership. Gate your pricing, your product specs, your competitive intel. You already do this for human visitors. Do it for bots too.

**c. Build for the narrow funnel.** AI answers can still drive traffic for specific types of queries. Unique data. Contrarian takes. Content an AI cannot plausibly generate on its own. Build for the narrow funnel, not the wide one. The PipeRocket data shows AI-referred visitors convert differently. They show up with less brand awareness (only 11.8% brand-name search intent versus 28.1% for organic), but they are further down the funnel. Separate research from mid-2026 found that AI search traffic converts at 14.2% compared to Google organic at 2.8%, a 5x advantage. The volume is lower but the quality is higher. Optimize for that.

**d. Watch the licensing market.** Over 50 publisher-AI agreements have been signed since 2023. Cloudflare is rolling out a Monetization Gateway for per-crawl payments. The infrastructure for getting paid is being built. You do not need to be the New York Times to participate. But you do need to be ready.

## 6. The Competitive Angle

Here is the question nobody wants to answer out loud. If you block ClaudeBot and your biggest competitor does not, who wins?

The answer depends on whether your buyers find you by name or by category.

A senior marketing leader researching "enterprise SEO platform" is searching by category. They do not have a brand in mind. If you are the only B2B brand visible in AI answers for that query, you own that discovery surface. If you blocked AI crawlers and your top three competitors did not, those three get every AI-generated comparison and recommendation. You are invisible. And the person who would have been your customer never knew you existed.

A CTO searching for "Atomic Glue case study" is searching by brand. They already know you. Blocking AI crawlers does not hurt that search. They will find your site through Google or direct navigation. The Rutger/Wharton data supports this: mid-sized publishers, which typically have stronger brand-direct traffic than category-discovery traffic, actually saw traffic increases after blocking AI crawlers. They were not losing category-level exposure because they did not have much to begin with, and they stopped AI from synthesizing their comparison content.

For B2B SaaS, the PipeRocket data is instructive. AI referrals make up under 9% of total traffic. But that under-9% pool converts differently. AI-referred visitors arrive with lower brand awareness (11.8% brand-name search intent versus 28.1% for organic), meaning they are discovery traffic, not brand-direct. They are the category searchers. If you block AI crawlers and your competitor stays open, that entire category-discovery pool goes to your competitor alone. For now, it is small. But AI search traffic converts at 14.2% compared to organic's 2.8%. Every percentage point of AI referral growth compounds at a premium conversion rate.

The reverse is also true. If you stay open and your competitors lock down, you own the entire AI discovery surface for your category. That is a meaningful moat in a world where the second-busiest crawler on the web is now an AI assistant, not a search engine.

**The zero-sum dynamic.** AI visibility is a fixed pie in a way that search visibility was not. Google could send a dozen results to a query and everyone got traffic. AI answers typically cite one or two sources. If you are not in that answer, you do not exist in that query. The margin between being cited and being invisible is razor-thin, and every publisher who blocks increases the citation share for those who stay open.

The early blocking data illustrates the divergence. Major publishers like CNN and the NYT, brands so big that name searches dwarf category searches, lost 23% of total traffic when they blocked. Mid-sized B2B publishers with 50,000 to 500,000 monthly visits saw no effect or traffic gains.

If you are a brand that people search for by name, blocking costs you more than it saves. If you are a brand that lives on category discovery, blocking costs you the entire AI channel.

Neither position is obviously right. But the strategic choice needs to be deliberate, not accidental. And it needs to be revisited quarterly, because the competitive landscape is shifting faster than any 12-month content plan can keep up with.

## 7. The Content Strategy Decision Framework

Here is the framework we use at Atomic Glue when a client asks what to do about AI bots. It applies to every piece of content you publish.

**Content Type x Bot Tier x Action**

**The crawl-to-referral audit.** Pull your server logs for the last 90 days. Identify every AI crawler that hit your site. Compare crawl volume to referral traffic from each source. If a bot is crawling thousands of pages and sending zero referred visitors, that is a relationship you should reconsider.

**The quarterly bot review.** Add a standing quarterly agenda item. Review which bots are crawling, what they are consuming, and what traffic they are sending back. Adjust your robots.txt and your content strategy accordingly. This should take 30 minutes. If your agency cannot produce this data, ask why.

## 8. The Timeline

Here is what we see happening and when it matters for your decision.

**Now (mid-2026).** Cloudflare has shifted the default to block AI crawlers for all new domains. Over 27% of websites are accidentally blocking AI crawlers without knowing it. Most B2B sites are running on default settings and have no idea what their bot traffic looks like. The blocking rate among top news publishers has climbed above 60%, versus under 10% among top retail domains. AI crawlers account for 20.3% of verified bot traffic and 57.5% of all HTML requests are now automated. The window for making a deliberate choice is closing faster than most marketing leaders realize.

**6 months.** Google's mixed-use crawler problem comes to a head. Publishers are pushing for separation between search crawlers and AI training crawlers. If Google splits them, the decision framework gets cleaner. If not, the ambiguity persists. Meanwhile, the crawl-to-refer ratios on every major AI platform are worsening as training volumes grow faster than search volumes.

**12 months.** The licensing market matures. Pay-per-crawl models move from pilot to standard. The question shifts from "should I block AI bots?" to "what is the right price for my content?" Publishers who can produce proprietary, high-demand data will negotiate from strength. Publishers running default settings will accept whatever terms are offered.

**18 months.** Agentic Internet becomes the dominant paradigm. Over 57% of HTML traffic is non-human today. That number climbs. The brands that built their AI visibility strategy early will have a durable competitive advantage. The ones that waited will find their category keywords already claimed by incumbents who made the deliberate choice while the window was open.

## 9. Five Questions Every B2B Leader Should Ask Their Agency

You should not need to be the expert on this. Your agency should be. Here are the questions that separate the ones who understand this from the ones who are still running last year's playbook.

**1. What AI crawlers are hitting my site right now, what are they consuming, and how much traffic are they sending back?**

If they cannot answer this with data from your actual server logs, they are flying blind. Not SimilarWeb estimates. Not Cloudflare dashboard screenshots. Your server logs. This is not a theoretical question. It is a reporting question.

**2. What is my crawl-to-referral ratio by bot?**

Not by source. By individual bot. Googlebot has a ratio of 5:1. GPTBot is 848:1. ClaudeBot is 4,580:1. Perplexity is 186:1. Treating them as one category means you are making policy decisions on bad information.

**3. What content on my site is most exposed to AI substitution?**

The answer is probably your comparison pages, your glossary content, and your "what is X" explainers. Those are the pages AI loves to synthesize. Your agency should know which pages those are before you ask. They should also know which pages drive your brand-name traffic versus your category-discovery traffic, because the blocking calculus is different for each.

**4. If we block training crawlers but keep discovery crawlers, what changes in our visibility?**

This is the nuanced question. Most AI companies separate these. Google does not. The Rutger/Wharton study found that the traffic loss from blocking was driven primarily by reduced brand exposure, not lost referral clicks. A good agency partner knows the difference and can model the impact of each approach on your specific content mix.

**5. What is our competitive position on AI visibility?**

If your top three competitors are visible in AI answers for your category keywords and you are not, that is a problem. If they have all blocked and you are the only brand showing up, that is an opportunity. Your agency should be monitoring this monthly, not annually.

## 10. What You Should Do Now

This is a marketing decision, not an IT decision. Own it.

Your content strategy needs a bot tier, just like it has audience tiers. Ask two questions about every piece of content:

Start measuring your crawl-to-referral ratio. Not just your pageviews. Track which bots are eating what. If a bot is devouring your most expensive content and sending nothing back, that should be a deliberate strategic choice. Not a default setting you forgot to change.

Add the quarterly bot review to your calendar. It takes 30 minutes. The cost of not doing it is making a strategic decision by accident.

## 11. The Window Is Open

We are in a transition period. The agentic internet is here. The business models are forming in real time.

Google's mixed-use crawler makes it hard to opt into search without also opting into AI. Cloudflare has declared a default block. Major publishers who blocked lost 23% of their traffic. Some mid-sized sites gained. The answer is not the same for everyone.

The brands that navigate this well will treat it as a relationship strategy, not a technical toggle. They will make deliberate choices about what to share, what to gate, and what to charge for. The ones that wait will find the choice has been made for them. And they will not like what they wake up to.

*Atomic Glue helps companies build, optimize, and scale high-performing websites, applications, and digital marketing systems. We work with marketing teams, startups, and agencies across the United States. If you want to know what your crawl-to-referral ratio looks like, we can show you in about an hour.*


Published July 20, 2026. Permalink: atomicglue.co/blog/your-website-is-their-training-data

Schedule a call

30 min · Video call

1
Date
2
Time
3
Details