Key takeaways

  • The decision is not binary. Training crawlers (GPTBot, ClaudeBot, Google-Extended) feed model training and can be blocked with zero search cost. Search and answer crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are what get you cited in AI answers, so blocking them removes you from those answers.
  • Blocking Google-Extended does not harm Google Search. Google states verbatim that it does not impact a site's inclusion in Search or act as a ranking signal, so you can opt out of Gemini training and keep full Search visibility and AI Overviews eligibility.
  • robots.txt is a request, not enforcement. Cloudflare caught Perplexity using a stealth crawler that spoofed Chrome and rotated IPs to fetch content from sites that had blocked it. Hard blocking needs WAF or edge rules, not robots.txt alone.
  • The traffic loss is real and worsening. Pew found users clicked a result link on only 8 percent of visits when an AI summary appeared, versus 15 percent without, and Ahrefs measured a 58 percent drop in position-one CTR on AI Overview keywords by December 2025.
  • You cannot judge block-versus-cite without seeing your AI traffic, and GA4 scatters most of it across Direct and Unassigned. A custom channel group is the only way to see the real picture.

Short answer: for most ad-funded publishers, block the AI crawlers that only train models if you object to that use, because you can do so at no cost to search visibility, but keep the search and answer crawlers open so you stay citable in AI answers. The block-versus-cite question is not binary. AI crawlers split into distinct types, and the type determines the consequence. Blocking a training crawler such as Google-Extended costs you nothing in Google Search. Blocking an answer crawler such as OAI-SearchBot removes you from ChatGPT's answers. And because robots.txt is voluntary, any block you genuinely want enforced has to happen at the edge, not in a text file that a determined crawler can ignore.

Should publishers block AI crawlers or allow them?

The mistake most publishers make is treating "AI crawlers" as one thing. They are not. They fall into three groups, and each has a different effect on your traffic and revenue.

Training crawlers exist to collect content that feeds future model training. OpenAI's GPTBot, Anthropic's ClaudeBot and Google's Google-Extended are the main ones. Blocking these has no effect on whether you appear in search results or AI answers. If you object to your work training commercial models, this is the group to block, and the cost is close to zero.

Search and answer crawlers are what put you inside AI answers. OpenAI's OAI-SearchBot surfaces your pages in ChatGPT search results. Anthropic's Claude-SearchBot improves Claude's search results. PerplexityBot feeds Perplexity's answers. Block these and you remove yourself from those answer engines. This is the group you almost certainly want to keep open.

User-triggered fetchers retrieve a page because a specific person asked for it in real time. OpenAI's ChatGPT-User, Anthropic's Claude-User and Perplexity-User sit here. Notably, ChatGPT-User does not obey robots.txt, because a live user-initiated fetch is treated as a user action rather than a crawl.

So "block or allow" is the wrong frame. The real question is which crawlers, doing which job, you want to permit. Here is the map.

CrawlerOperatorJobObeys robots.txtBlocking effect
GPTBotOpenAIModel trainingYesNone on search or citation
OAI-SearchBotOpenAIChatGPT search citationsYesRemoved from ChatGPT answers
ChatGPT-UserOpenAILive user fetchNoCannot be blocked via robots.txt
ClaudeBotAnthropicModel trainingYesNone on search or citation
Claude-SearchBotAnthropicClaude search resultsYesWeakens Claude citations
Claude-UserAnthropicLive user fetchYesBlocks user-directed answers
Google-ExtendedGoogleGemini training and groundingYesNone on Google Search
PerplexityBotPerplexityPerplexity answersNominallyRemoved from Perplexity (if respected)

What do you lose by blocking AI from your content?

Blocking has two very different price tags depending on which crawlers you block, and it is worth being precise, because the wrong block quietly costs you an emerging revenue channel.

If you block only the training crawlers, you lose almost nothing measurable. Google is explicit on this point. Its documentation states that Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. Google-Extended governs only whether your content trains future Gemini models and grounds Gemini and Vertex AI. The crawler that actually indexes and ranks your pages is Googlebot, which is entirely separate. So you can opt out of Gemini training and keep full Search rankings and AI Overviews eligibility. The same logic applies to GPTBot and ClaudeBot: they train models, they do not feed answer citations, and blocking them does not remove you from ChatGPT search or Claude search.

If you block the search and answer crawlers, the cost is real and growing. AI referral traffic is still small in absolute terms, but the evidence points to it converting far better than ordinary organic search. A Semrush study published in June 2025 found that the average visitor arriving from an LLM was worth roughly 4.4 times an organic search visitor, with ChatGPT referrals converting at 15.9 percent against 1.76 percent for Google organic, while AI referral traffic, then about 1 percent of all referrals, was growing 527 percent year on year. These are high-intent visitors arriving with a question already half-answered. Block the answer crawlers and you forfeit that channel entirely, at the exact moment it is scaling. This is the core tension: the crawlers that cost you nothing to block are not the ones taking your search traffic, and the ones that could earn you traffic are the ones a blanket block would shut off.

One honest caveat: the conversion advantage is not universal. Studies of large ecommerce samples have found AI referrals converting no better, or worse, than other channels, so the size of the prize varies by vertical and intent. That is precisely why the decision has to come from your own numbers rather than a headline, a point we return to below.

When does blocking AI make sense for a site?

Blocking makes sense in specific, defensible situations rather than as a default reflex.

Block training crawlers when you object on principle or licensing grounds to your work training commercial models. Because the search cost is zero, this is a low-risk stance. You can block GPTBot, ClaudeBot and Google-Extended and lose no ranking, no AI Overview eligibility and no answer-engine presence.

Consider blocking answer crawlers only when AI answers are demonstrably substituting for your traffic rather than sending any, and you have the data to prove it. If you are being cited heavily but seeing no referral clicks and no conversions, the citation is pure extraction with no return, and blocking becomes rational. But confirm this from your own numbers first, because the general trend runs the other way for many sites.

Understand why the pressure to block is rising. The traffic loss driving these decisions is well documented. Pew Research Center's browsing-panel study of 900 US adults across 68,879 unique searches found that users clicked a traditional search-result link on just 8 percent of visits when an AI summary appeared, versus 15 percent when it did not, and clicked one of the AI summary's own cited sources on only 1 percent of visits. The same study found users ended their session entirely on 26 percent of pages showing an AI summary, against 16 percent for standard results. On the search-engine side, Ahrefs analysed 300,000 keywords using aggregated Search Console data and found that AI Overviews cut position-one organic CTR by 58 percent as of December 2025, up from a 34.5 percent reduction in April 2025. The trend is worsening, which is why the block question is on the table at all. But note the asymmetry: this damage comes from Google's AI Overviews, and you cannot block your way out of it, because the crawler behind AI Overviews is Googlebot, the same crawler you need for ordinary rankings.

Can you allow answer engines but block training?

Yes, and for most publishers this is the sensible middle path. Because each crawler has its own user-agent token and its own robots.txt directive, you can permit citation while refusing training. A configuration that blocks training and keeps you citable looks like this.

  • Block training only: GPTBot, ClaudeBot, Google-Extended (Disallow: /)
  • Keep open for citation: OAI-SearchBot, Claude-SearchBot, PerplexityBot
  • Leave user fetchers alone: ChatGPT-User (ignores robots.txt anyway), Claude-User, Perplexity-User

There is one important caveat that turns this from a clean answer into a two-layer problem: robots.txt is voluntary. Compliant bots honour it. Non-compliant ones do not. In August 2025 Cloudflare documented Perplexity running a stealth, undeclared crawler that spoofed a Chrome-on-macOS user agent, used IPs outside its published range and switched ASNs to fetch content from sites that had explicitly blocked PerplexityBot, generating 3 to 6 million stealth requests a day on top of 20 to 25 million from its declared crawler. Cloudflare de-listed Perplexity as a verified bot in response. The lesson is blunt: a robots.txt block stops the well-behaved crawlers and does nothing to the determined ones.

Anthropic makes the same point about a different mechanism. It publishes crawler IP ranges but warns that IP-based blocking is unreliable, because blocking the IPs can stop a bot from reading your robots.txt opt-out in the first place. So robots.txt is the right signal for compliant bots, and IP or WAF blocking is the enforcement layer for the rest. If you actually need a block enforced, do it at the edge with WAF rules or a challenge, and treat robots.txt as the polite front door rather than the lock.

Two layers, then:

  1. robots.txt for the opt-out signal that compliant crawlers respect (training-only blocks, citation kept open).
  2. WAF or edge rules for anything you need genuinely enforced against crawlers that spoof user agents and rotate IPs.

How do I decide based on my own data?

Every recommendation above collapses into one requirement: you cannot make this call sensibly without seeing what AI is doing to your specific site. And by default you cannot see it, because GA4 hides most AI referral traffic.

Referrals from chatgpt.com, perplexity.ai, claude.ai, gemini.google.com and copilot.microsoft.com often arrive with no UTM tags and no referrer header, because AI apps and embedded browsers strip it. GA4 therefore scatters them across Referral, Direct and Unassigned. Google only added a native "AI Assistant" default channel on 13 May 2026, and even that undercounts badly: analyses of AI referral traffic have found roughly 35 to 70 percent of AI sessions arriving without referrer data and landing in Direct. The native channel is a floor, not a full count. To see the real picture you need a Custom Channel Group with a source or referrer regex, ordered above the broad Referral rule so AI sessions match the AI channel first. A regex covering the main engines looks like this: chatgpt\.com|openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com|you\.com.

Once you can actually see AI traffic, the decision becomes a worked comparison rather than a guess. Ask three questions of your own data.

  • Is AI referral traffic arriving at all, and is it converting? If AI-referred visitors are converting well above your organic baseline, staying citable is clearly worth it and a blanket answer-crawler block would be self-harm.
  • Is AI substituting for search clicks you used to get? Cross-reference the AI Overview CTR decay in Search Console against your own click trend. If impressions hold but clicks fall on AI Overview queries, that is the extraction pattern.
  • What does AI-referred traffic earn versus your site average? For an ad-funded publisher this is the number that settles the argument. If AI sessions monetise at or above your average RPM, they are a channel to protect, not block.

Search Console can help with the search side. Google launched dedicated generative-AI performance reports for AI Overviews and AI Mode on 3 June 2026, rolling out to a subset of properties first. Read the numbers cautiously: when a URL appears both inside an AI Overview and as a blue link on the same results page, Search Console counts it as a single impression, and a separate logging error over-reported impressions from 13 May 2025 until Google fixed it in late April 2026 (clicks were unaffected). Treat AI-surface impression trends as directional, not exact.

If building a custom GA4 channel group, wiring the regex and reconciling Search Console AI reports sounds like a project, that is precisely the gap Ramprt closes. It reads your existing GA4 read-only, with no new tracking script, and shows AI referrals by engine, which AI crawlers are taking your content, how often AI answers are built on your pages, and what AI-referred traffic earns against your site average. You can see the whole picture in the live demo before deciding anything. Deciding block-versus-cite without that view is guessing.

The bottom line

Block training crawlers if you object to your work training models, because it costs you nothing in search. Keep the search and answer crawlers open so you stay citable, because that channel is small but high-converting and growing fast. Enforce any hard block at the WAF or edge rather than trusting robots.txt, because compliant bots obey it and the ones worth worrying about do not. And make all of these calls from your own traffic data, not from headlines, because the right answer for your site is written in numbers GA4 does not show you by default.

For the enforcement mechanics, see our guide on robots.txt versus WAF blocking, and for the measurement side, how to track AI traffic in GA4.

Frequently asked questions

Does blocking Google-Extended hurt my Google Search rankings?

No. Google's own documentation states that Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal. Google-Extended only controls whether your content trains future Gemini models and grounds Gemini and Vertex AI. The crawler that indexes and ranks your pages is Googlebot, which is separate, so you can block Google-Extended and keep full Search visibility and AI Overviews eligibility.

If I block AI training crawlers, do I disappear from ChatGPT and Perplexity answers?

Not if you block only the training crawlers. GPTBot and ClaudeBot feed model training, not answer citations. The crawlers that put you in AI answers are OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude, and PerplexityBot for Perplexity. Keep those open and you stay citable while still blocking training.

Is robots.txt enough to keep AI out of my content?

No. robots.txt is a request, not enforcement. Compliant crawlers honour it, but in August 2025 Cloudflare documented Perplexity using a stealth crawler that spoofed a Chrome user agent and rotated IPs to fetch content from sites that had blocked it. For a block you genuinely need enforced, use WAF or edge rules rather than robots.txt alone.

Why can't I see AI referral traffic in GA4?

AI referrals from ChatGPT, Perplexity, Claude and others often arrive with no UTM tags and no referrer, because AI apps strip it, so GA4 scatters them across Referral, Direct and Unassigned. Google added a native AI Assistant channel on 13 May 2026, but roughly 35 to 70 percent of AI sessions still land in Direct. To see the real figures you need a custom channel group with a regex on AI hostnames ordered above the broad Referral rule.

Is AI referral traffic worth keeping if it is so small?

For many publishers, yes. It is small in volume but high in intent. A Semrush study found the average LLM visitor worth about 4.4 times an organic search visitor, with ChatGPT referrals converting at 15.9 percent versus 1.76 percent for Google organic, and AI referral traffic growing 527 percent year on year. The advantage is not universal and varies by vertical, so confirm it against your own data before deciding, but for most sites the middle path of blocking training while staying citable makes sense.