Key takeaways
- To stay citable, allow the search and user agents (OAI-SearchBot, Claude-SearchBot, PerplexityBot and their user bots) even if you block the training bots. Blocking GPTBot, ClaudeBot or Google-Extended stops training, not citation.
- Mentions beat links. Ahrefs' 75,000-brand study found branded web mentions correlate with AI visibility at 0.66 to 0.71, and YouTube mentions at about 0.737, while the number of backlinks shows very weak correlation.
- The content tactics with the strongest evidence are statistics, direct quotations from named authorities, and cited sources. The peer-reviewed GEO study found these lifted visibility in generative engines by up to 40 percent.
- Two popular tactics have almost no evidence: llms.txt (97 percent of valid files got zero requests in Ahrefs' 137,000-site study) and generic schema markup (no measured lift in AI Overview citations).
- Citation is often the only visibility you get. Pew found users clicked a link on just 8 percent of visits when an AI summary appeared, versus 15 percent without, so measure net clicks, not citation count.
The short answer: to get cited by AI, first make sure the right AI crawlers can reach and index your pages, then earn brand mentions across the web and give AI engines the raw material they prefer to quote. Allow the search and user-facing bots (OpenAI's OAI-SearchBot, Anthropic's Claude-SearchBot, Perplexity's PerplexityBot and their user agents) in robots.txt even if you block the training bots. Lead each page with a direct answer, back it with specific statistics, quote named sources, and get talked about on the wider web and on YouTube. Two things you can safely skip: llms.txt and adding schema markup purely to win AI citations. Neither has evidence behind it.
How do I get cited by AI like ChatGPT and Perplexity?
Getting cited breaks down into two jobs: being reachable by the right crawlers, and being the passage an engine wants to quote. Most publishers get the first job wrong without realising it, because the AI bot landscape is more granular than people assume.
Each of the four major AI companies runs several separately named crawlers, and you can control each one independently in robots.txt. The critical distinction is between training bots and search or citation bots.
| Company | Training crawler | Search / citation crawlers |
|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot (ChatGPT search citations), ChatGPT-User (user-initiated fetch) |
| Anthropic | ClaudeBot | Claude-SearchBot (search indexing), Claude-User (user-initiated fetch) |
| Perplexity | n/a | PerplexityBot (indexing and linking), Perplexity-User (live browsing) |
| Google-Extended (a robots token, not a crawler) | Googlebot (the live search index AI Overviews draw on) |
Here is the trap. Blocking GPTBot, ClaudeBot, Google-Extended or CCBot blocks training, not real-time citation. ChatGPT search, Claude search and Perplexity cite pages through their SearchBot and User agents. Google AI Overviews draw on the live Google Search index via Googlebot, not Google-Extended. Google itself confirms that blocking Google-Extended stops Gemini training but does not remove you from AI Overviews. So over-blocking training bots does not protect you from being summarised. It only removes you from the training corpus while leaving you visible in AI answers, now without any say over future training.
If your goal is to be cited while opting out of training, the robots.txt logic is: allow the search and user agents, disallow the training agents if you wish.
| Allow (to stay citable) | Disallow (to opt out of training) |
|---|---|
| OAI-SearchBot, ChatGPT-User | GPTBot |
| Claude-SearchBot, Claude-User | ClaudeBot |
| PerplexityBot, Perplexity-User | CCBot (Common Crawl) |
| Googlebot (never block this) | Google-Extended |
A few operational notes. Verify a bot by checking its IP against the operator's published list rather than trusting the user-agent string, which anyone can spoof. OpenAI and Perplexity both publish IP ranges for their crawlers for exactly this purpose. Anthropic does not publish ClaudeBot IP ranges and uses provider public IPs, so control it via robots.txt tokens, not by IP blocking, which can even stop it reading your robots.txt in the first place. After an OpenAI robots.txt change, allow roughly 24 hours for it to take effect.
Do I need high domain authority to be cited by AI?
No, and this is the most useful finding for smaller publishers. The strongest published correlate of AI visibility is being mentioned, not being linked.
Ahrefs studied 75,000 brands across ChatGPT, AI Mode and AI Overviews. Branded web mentions correlated with AI visibility at roughly 0.66 to 0.71, while the number of backlinks showed very weak correlation. YouTube mentions correlated most strongly of all, at about 0.737. The practical read is blunt: earn brand mentions and get talked about across the web and on YouTube, rather than chasing links alone. A brand that is discussed widely, even without a hyperlink, is more visible to AI than one sitting on a pile of backlinks.
That said, crawlability and indexing remain the single most important prerequisite. Pages blocked by robots.txt, hidden behind a hard paywall, or returning errors are effectively invisible to AI. And ranking still helps, it is just no longer sufficient. Ahrefs found that 38 percent of AI Overview citations come from a page that also ranks in the top 10, down from about 76 percent in July 2025. The drop is explained by Google's "query fan-out": it splits a query into related sub-queries and cites the pages that appear most often across those sub-query SERPs, not just the original query. So a page ranking modestly for a specific sub-question can be cited even if it does not rank for the head term.
Put together: you do not need high domain authority, but you do need to be crawlable, indexed, mentioned, and specific enough to win the sub-queries.
What kind of content do AI engines cite most?
This is where the evidence is clearest. The peer-reviewed GEO study by Aggarwal et al. (KDD 2024) tested content changes against real generative engines and found that methods such as adding statistics, direct quotations from authorities, and cited sources boosted visibility by up to 40 percent. Keyword stuffing gave little to no improvement.
Translated into a working checklist for a publisher:
- Answer first. Open each section with a direct, self-contained answer to a specific question. AI engines extract clean passages, and the query fan-out rewards pages that resolve narrow sub-questions.
- Include original data and specific numbers. Statistics were one of the top-lifting tactics. Unique data you publish yourself is harder for anyone else to be cited for.
- Quote and cite named sources. Direct quotations from authorities and cited sources both lifted visibility. Name the source and link it.
- Write extractable passages. Short, standalone paragraphs, clear tables and worked examples give the engine a clean chunk to lift.
There is corroborating commercial evidence that citation is worth the effort. Seer Interactive, analysing 53 brands across 5.47 million queries and 2.43 billion organic impressions, found that brand-cited pages get about 120 percent more organic clicks per impression than uncited pages on AI Overview SERPs. Being cited does not just protect visibility, it appears to lift the clicks you still get, even though cited pages still trail a pre-AI-Overview baseline.
Worked example: a citable passage versus a non-citable one
Non-citable: "Bank holidays can affect trading, and there are various things to keep in mind depending on where you are and what markets you follow."
Citable: "UK markets close on eight bank holidays in 2026. The London Stock Exchange is shut on those dates, so settlement shifts to the next business day. According to [named authority], holiday-thinned volume widens spreads by [figure]." The second version leads with a direct answer, contains specific numbers, and quotes a named source. That is the shape AI engines quote.
What tactics do NOT work for AI citation?
Two popular tactics have little to no evidence behind them, and it is worth not wasting time on either.
llms.txt. This proposed file is not used by any major AI service. Google's John Mueller has called it "purely speculative for now", and an Ahrefs analysis of 137,000 sites found that, of the roughly 38,000 with a valid llms.txt, 97 percent received zero requests for the file in the month studied. Among the handful that did see traffic, most requests came from non-AI bots. If nobody is fetching it, it cannot be a citation lever.
Schema and structured data (for AI citation specifically). Schema did not increase AI Overview citations in controlled tests, and there are no peer-reviewed studies showing schema drives LLM citations. Schema still earns you classic rich results and improves machine-readability, so keep it for those reasons, but do not add it expecting it to win AI citations on its own.
Keyword stuffing. The GEO study found it offered little to no improvement.
The honest summary: spend your effort on mentions, statistics, quotations and cited sources, not on llms.txt or on schema-as-a-citation-hack.
How do I check if AI is citing my pages?
This is harder than it should be, because AI-referred traffic is deliberately murky in standard analytics. You need to know where each type of AI exposure lands.
| AI source | Where it shows up | The catch |
|---|---|---|
| ChatGPT, Perplexity, Gemini, Copilot, Claude clicks | GA4 Referral channel | Only when a referrer is actually sent |
| Google AI Overviews / AI Mode clicks | GA4 Organic Search | Often no referrer, so they cannot be isolated as "AI" |
| AI Overviews / AI Mode in Search Console | Standard Web performance report | Folded in with normal search, cannot be cleanly separated |
| Generative AI features report (June 2026) | Dedicated Search Console view | Impressions only: no clicks, CTR, position or query data |
Two things follow. First, to see AI referrals in GA4 at all, you need a custom channel group or regex over referrer domains, and even then it is a partial view because much AI exposure sends no click. Second, in Search Console the new dedicated Generative AI features report, launched June 2026, shows impressions only. An impression there means your URL appeared inside an AI feature, either as a citation in an AI Overview or as a referenced source in AI Mode. There are no clicks, CTR, position or query data, and a URL shown in both an AI feature and the blue links on one search counts as a single impression.
Because the picture is fragmented, the metric that matters is net clicks over time, not raw citation count. Citation is a means, clicks and revenue are the end. If you want a single clean read of what AI is doing to your existing GA4, without adding another tracking script, you can try Ramprt free at ramprt.io/demo. It reads your GA4 read-only and its AI tab isolates AI referrals by engine, which crawlers are taking your content, and what AI-referred traffic earns versus your site average.
How long does it take to start getting cited by AI?
The research does not give a single number, so here is the mechanism rather than a fabricated figure. Two clocks are running.
The first is technical and fast. After you change robots.txt to allow the search and user agents, OpenAI's guidance is that changes take roughly 24 hours to take effect. Once a crawler can reach a page and it is indexed, it becomes eligible for citation on the next relevant query.
The second clock is slow and is the one that actually governs your visibility: earning mentions. Because branded mentions and YouTube mentions are the strongest correlates of AI visibility, and those accumulate through PR, coverage and being talked about, this is a months-not-days effort. The content-level changes (answer-first passages, statistics, quotations, cited sources) sit in between: they take effect as pages are re-crawled and re-indexed, so weeks rather than months for a site that is already crawled regularly.
Sequence it like this: fix crawlability today (fast win), rewrite key pages for extractability over the coming weeks, and invest in mentions over the coming quarters. Then judge it all by net clicks, not by how many citations you can spot.
Why does getting cited matter so much now?
Because the citation is increasingly the only visibility you get. Pew Research studied the browsing of 900 US adults across 68,879 Google searches in March 2025, of which about 18 percent produced an AI summary. When an AI summary was shown, users clicked a traditional link on only 8 percent of visits, versus 15 percent without a summary, and clicked a source link inside the summary on just 1 percent of visits.
That is the whole case in one statistic. AI summaries roughly halve onward clicks, so if your page is not the one cited inside the summary, you are largely invisible on that search. Getting cited becomes a goal in its own right. And because the click erosion is real, the honest way to run this is to measure whether citation is translating into net clicks and revenue, not to celebrate a citation count that may not pay a penny.
A practical order of operations
- Audit robots.txt. Allow OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. Never block Googlebot. Block training bots (GPTBot, ClaudeBot, Google-Extended, CCBot) only if you actively want to opt out of training.
- Confirm your key pages are crawlable, indexed, and not stuck behind a hard paywall or throwing errors.
- Rewrite priority pages answer-first, with specific statistics, named quotations and cited sources. Target specific sub-questions, not just head terms.
- Invest in brand mentions across the web and YouTube. This is the strongest lever and the slowest.
- Do not build llms.txt or add schema expecting AI citations from it.
- Set up AI referral tracking in GA4 (custom channel group), watch the Search Console Generative AI features report for impressions, and judge success by net clicks over time.
For related reading, see our guides on what AI is doing to your traffic and how to track AI referrals in GA4.
Frequently asked questions
Should I block GPTBot and ClaudeBot?
Only if you specifically want to opt out of AI model training. Blocking GPTBot and ClaudeBot stops training, not citation. To stay citable while opting out of training, allow the search and user agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User) and block only the training bots. Never block Googlebot, since Google AI Overviews draw on the live Google Search index.
Does blocking Google-Extended remove me from AI Overviews?
No. Google confirms that blocking Google-Extended stops Gemini training but does not remove you from AI Overviews, which draw on the live Googlebot search index. Google-Extended is a robots.txt token, not a separate crawler, and it only governs training use.
Do backlinks help me get cited by AI?
Only weakly. Ahrefs' 75,000-brand study found the number of backlinks shows very weak correlation with AI visibility, while branded web mentions correlate at 0.66 to 0.71 and YouTube mentions at about 0.737. Earning mentions and being talked about matters more than link count.
Is llms.txt worth setting up?
No. No major AI service fetches it. Google's John Mueller has called it purely speculative, and Ahrefs' analysis of 137,000 sites found 97 percent of domains with a valid llms.txt received zero requests for the file. It is not a citation lever.
Can I see AI citations in Google Search Console?
Partly. AI Overviews and AI Mode data is folded into the standard Web performance report. The dedicated Generative AI features report, launched June 2026, shows impressions only, with no clicks, CTR, position or query data, and counts a URL shown in both an AI feature and blue links as a single impression.
Sources
- Pew Research Center: Google users are less likely to click on links when an AI summary appears
- Ahrefs: An analysis of AI brand visibility factors (75K brands)
- Ahrefs: 38% of AI Overview citations pull from the top 10
- GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024)
- Ahrefs: We analysed 137K sites, 97% of llms.txt files never get read
- Search Engine Journal: Google says llms.txt is purely speculative for now
- Google Search Central: Introducing Search Generative AI performance reports
- Seer Interactive: AIO impact on Google CTR (2026 update)