# AI Crawlers: Which Ones Get You Cited?
> Seven weeks of first-party logs, 761,885 hits. Most AI crawlers cannot cite you at all. Here is which ones can, and what to do about the rest.

Canonical URL: https://cite.solutions/blog/ai-crawlers-which-ones-get-you-cited
Source: Cite Solutions (cite.solutions)
Published: 2026-07-28
---

[Research](/category/research)11 min read

# AI Crawlers: Which Ones Get You Cited?

[Subia PeerzadaFounder, Cite Solutions · July 28, 2026](https://www.linkedin.com/in/subia-peerzada-75025764/)

Key takeaways

## Key takeaways for AI citation readiness

Make every important page easier for answer engines to quote, trust, and reuse.

1. 01Lead each section with a direct answer block before expanding into detail.
2. 02Put evidence close to the claim so AI systems can extract support cleanly.
3. 03Use schema and strong information architecture to improve eligibility, not as a gimmick.

We run edge middleware on cite.solutions that writes a row every time an AI bot touches the site. Seven weeks in, it has logged 761,885 requests from AI crawlers.

The headline number went up almost tenfold over that period. It would look great in a board deck.

It means almost nothing. Once you split those hits by what the bot was actually doing, the only category that can put your brand in an answer turns out to be 3% of the traffic, and it is shrinking while everything else grows.

## Which AI crawlers get you cited?

Only answer-time crawlers can get you cited: ChatGPT-User, Claude-User, PerplexityBot, and OAI-SearchBot. Training crawlers like GPTBot, ClaudeBot, and Meta-ExternalAgent read your content for model training and never link back. In our own logs, answer-time fetches were 3% of 761,885 AI crawler hits over seven weeks.

First-party crawler logs

Only 3% of AI crawler traffic can ever cite you

Every AI bot hit on cite.solutions across seven weeks, June 9 to July 27, 2026, split by what the fetch was actually for.

22,955

answer-time fetches out of 761,885 total AI crawler hits

The other 738,930 requests read the site without any path to naming it in an answer.

Training crawls

626,271 · 82.2%

Content taken for model training. Cannot produce a citation.

General and index

112,659 · 14.8%

Builds the candidate pool. Cites you only indirectly, later.

Answer-time fetches

22,955 · 3%

A person asked a question and the engine came to read us.

Answer-time fetches per week

Total crawl volume rose almost tenfold over the same seven weeks. The only traffic that can cite us fell 77% from its peak.

4,604

4,045

5,996

3,283

2,365

1,296

1,366

Jun 15

23,224

Jun 22

34,521

Jun 29

201,161

Jul 6

64,086

Jul 13

59,194

Jul 20

151,164

Jul 27

228,535

Top row: answer-time fetches. Bottom row: all AI bot hits that week.

Source: Cite Solutions first-party edge-middleware bot log, cite.solutions, 7 weeks ending 2026-07-27, 761,885 classified AI bot requests

## AI crawlers do three different jobs and only one of them cites you

Every vendor runs a small fleet, not a single bot. The user agents look interchangeable in a log file. They are not, and the difference decides whether a fetch can ever turn into a citation.

### Training crawlers take your content and give nothing back

GPTBot, ClaudeBot, Meta-ExternalAgent, Amazonbot, and Applebot-Extended read pages to build training corpora. The content gets tokenized, folded into a future model version, and surfaces months later as unattributed knowledge.

There is no link, no referral, and no citation. [Cloudflare's 2026 bot report](https://blog.cloudflare.com/agentic-internet-bot-report/) put training at 52% of all crawler requests as of June 2026, up from 22% in spring 2025.

> Most AI crawler traffic is not an audience. It is a download.

### Index crawlers build the candidate pool you might get picked from

OAI-SearchBot is the clearest example. [OpenAI's own crawler documentation](https://developers.openai.com/api/docs/bots) describes it as the bot "used to surface websites in search results in ChatGPT's search features." Block it and you drop out of the pool entirely.

An index crawl is necessary but not sufficient. Being in the pool is not being in the answer.

### Answer-time fetchers are the bots that can name you today

ChatGPT-User is the one that matters most. OpenAI documents it as the agent that visits pages "when users ask questions," rather than crawling on a schedule. Claude-User and Perplexity's user-triggered fetches work the same way.

When one of these hits your server, a real person asked a real question thirty seconds ago and the engine came to read you before writing its answer. That is the closest thing to a live citation signal any log file gives you.

**What most AI crawler reports count:**

* •Total bot hits per week
* •Number of distinct user agents seen
* •Whether GPTBot is allowed in robots.txt
* •Which vendor sent the most traffic

**What actually predicts a citation:**

* •Answer-time fetches only, broken out by user agent
* •Which specific URLs those fetches landed on
* •Whether the fetch returned a clean 200 with the answer in server HTML
* •Whether that number is rising or falling week over week

## What 761,885 AI crawler hits on our own site showed

Here is the full weekly series from our bot log. Every row is one week of classified requests to cite.solutions. This is a single mid-size B2B site, so treat it as one honest instrument reading rather than a market-wide census.

| Week ending | All AI bot hits | Training    | General / index | Answer-time | Answer-time share |
| ----------- | --------------- | ----------- | --------------- | ----------- | ----------------- |
| Jun 15      | 23,224          | 13,895      | 4,725           | 4,604       | 19.8%             |
| Jun 22      | 34,521          | 22,538      | 7,938           | 4,045       | 11.7%             |
| Jun 29      | 201,161         | 152,308     | 42,857          | 5,996       | 3.0%              |
| Jul 6       | 64,086          | 45,393      | 15,410          | 3,283       | 5.1%              |
| Jul 13      | 59,194          | 48,001      | 8,828           | 2,365       | 4.0%              |
| Jul 20      | 151,164         | 136,840     | 13,028          | 1,296       | 0.9%              |
| Jul 27      | 228,535         | 207,296     | 19,873          | 1,366       | 0.6%              |
| **Total**   | **761,885**     | **626,271** | **112,659**     | **22,955**  | **3.0%**          |

Five things fall out of that table.

### Finding #1: Only 3% of AI crawler traffic could ever produce a citation

Answer-time fetches came to 22,955 of 761,885 requests. The remaining 738,930 read the site with no mechanism to name it in an answer.

Our split is more extreme than Cloudflare's network-wide 52% training figure, and the reason is instructive: a single site with a small number of high-value pages attracts repeat training crawls far out of proportion to its size.

### Finding #2: Answer-time fetches fell 77% while total volume rose tenfold

The peak was 5,996 answer-time fetches in the week ending June 29\. Five weeks later it was 1,366, a 77% drop. Total AI bot hits across the full seven weeks went the other way, from 23,224 to 228,535.

Any dashboard reporting "AI crawler traffic" as one number would have shown a triumphant curve while the only meaningful line collapsed underneath it.

> Volume went up tenfold. The traffic that can actually cite us fell by three quarters.

### Finding #3: One training crawler produced 82% of our traffic and zero citations

In the week ending July 27, Meta-ExternalAgent alone accounted for 187,831 hits, 82.2% of everything. Meta documents it as a crawler for training Llama and Meta AI.

It was not discovering anything. It hit 10,836 unique paths at an average of 17.3 fetches each, put 81% of its volume on just 40 pages, and pulled `/contact` 6,115 times in a single week. That is a re-fetch loop.

Strip Meta out and the week was 40,704 hits against 39,753 the week before. Flat. The 51% growth in our headline number was one misbehaving bot.

### Finding #4: OpenAI is now 96% of the answer-time traffic we see

Of 1,366 answer-time fetches in the most recent week, 1,314 came from OpenAI user agents. ChatGPT-User contributed 760 and OAI-SearchBot 554.

For a site in our category, optimizing for answer-time retrieval currently means optimizing for one vendor. That concentration is a risk, and we have written about [why betting everything on ChatGPT is dangerous](/blog/ai-visibility-chatgpt-concentration-risk).

### Finding #5: Perplexity and Anthropic have effectively stopped fetching at answer time

PerplexityBot ran 2,187 answer-time fetches in the week ending June 29\. By July 6 it was 97\. It has sat between 21 and 29 every week since.

Anthropic's Claude-User has never cleared 25 in any week we have measured. Both engines still cite sources in their answers, so they are clearly retrieving from somewhere. They are just not retrieving from us at answer time, which is its own diagnosis.

### Do you know how many of your AI crawler hits could actually cite you?

We instrument your logs, separate answer-time fetches from training noise, and show you which pages the citing bots are reading. First findings inside 14 days.

[Book a Discovery Call](/contact)

## The AI crawler reference table: who each bot is and what to do with it

This is the working table we keep. The column that matters is the third one.

| User agent         | Operator   | Can it cite you?   | What it is actually doing                                          | Our default                                  |
| ------------------ | ---------- | ------------------ | ------------------------------------------------------------------ | -------------------------------------------- |
| ChatGPT-User       | OpenAI     | Yes, directly      | Fetches a page the moment a user asks a question                   | Never block. Highest-value bot on the site.  |
| OAI-SearchBot      | OpenAI     | Yes, via the index | Builds the pool ChatGPT search draws from                          | Allow. Keep money pages fast and clean.      |
| PerplexityBot      | Perplexity | Yes                | Indexes and fetches for cited answers                              | Allow, and watch the volume trend.           |
| Claude-User        | Anthropic  | Yes, in principle  | User-triggered fetch for Claude web search                         | Allow. Negligible volume for us so far.      |
| GPTBot             | OpenAI     | No                 | Training corpus for future models                                  | Allow, with no citation expectation.         |
| ClaudeBot          | Anthropic  | No                 | Training corpus for Claude                                         | Allow.                                       |
| Meta-ExternalAgent | Meta       | No                 | Training for Llama and Meta AI                                     | Rate-limit. 82% of our volume, zero return.  |
| Amazonbot          | Amazon     | No                 | Training and general collection                                    | Rate-limit if egress matters.                |
| Applebot-Extended  | Apple      | No                 | Training opt-in for Apple Intelligence                             | Policy call, not a visibility call.          |
| Google-Extended    | Google     | No                 | Not a crawler. A robots.txt token controlling Gemini training use. | Leave alone unless you want out of training. |

Two notes worth keeping straight. Google-Extended is a permission flag, not a bot, so looking for its hits in your logs will waste an afternoon. And blocking OAI-SearchBot removes you from ChatGPT search answers even though it never appears as a "citing" bot in your reports.

The distinction between OpenAI's index bot and its answer-time bot is the single most common configuration error we find. We covered the robots.txt mechanics in [is ChatGPT-User allowed in your robots.txt](/blog/chatgpt-user-robots-txt-ai-citations).

> The crawler that hits you most is usually the one that will never name you.

## Step 1: Split answer-time fetches out of your log before you report anything

Take your server or edge logs, classify every AI user agent into training, index, or answer-time, and report the three lines separately. One blended "AI bot traffic" number hides the only signal in the data.

If you do not have a persistent store, this is the first thing to build. Runtime logs on most hosts retain well under an hour, which is how teams end up with months of zeros and assume no bots are visiting.

## Step 2: Check which URLs the answer-time bots actually landed on

Filter to answer-time fetches only, then group by path. You are looking for whether the engines read the pages you want quoted or whether they keep landing on the homepage.

Ours is blunt about this. `/` takes most of the answer-time traffic and exactly one content page gets quoted with any consistency: [our AI search market share analysis](/blog/ai-search-market-share-2026), at 137 answer-time fetches in the most recent week. Everything else in a 290-post library is being read for training and never at answer time.

## Step 3: Rate-limit the training crawlers that only cost you money

A training crawl consumes origin bandwidth and returns nothing. When one bot generates 188,000 redundant fetches in a week, that is a hosting bill, not a marketing channel.

Set a crawl-delay in robots.txt for the offender, or block it at the edge if it ignores the directive. Do this by user agent, deliberately, and never to a bot in the "can cite you" column.

## Step 4: Fix what the answer-time bots read when they arrive

An answer-time fetch is a live audition and you get one pass. The engine reads what your server returns, not what renders after JavaScript loads.

Put a direct 40 to 60 word answer under every question-shaped heading, in server HTML. That is the passage extraction problem, and we broke it down in [passages beat pages](/blog/passages-beat-pages-how-to-structure-content-for-ai-citation). If your rendered and server HTML disagree, run an [HTML parity audit](/blog/html-parity-audit-ai-retrieval) before anything else.

> An answer-time fetch is a live audition. Your server HTML is the performance.

## Step 5: Track answer-time fetches weekly and ignore the headline total

Fix the classification, fix the reporting day, and chart one line: answer-time fetches per week, split by vendor. That number moving is the earliest leading indicator of a citation change you can get from your own infrastructure.

Total bot volume is noise driven by whichever training crawler is having a busy week. If you would rather not staff the instrumentation and the content work together, [a managed GEO agency can run both](/geo-agency), and an [AI visibility audit](/ai-visibility-audit) will tell you where you stand before you commit to the operations.

## What will not move this

Three habits keep showing up and none of them touches the mechanism.

Blocking training crawlers to "protect your content" does nothing for citations. It is a rights and cost decision, and a defensible one, but it will not get you named in an answer.

Publishing llms.txt so crawlers can find your best pages is a hope, not a lever. A 90-day audit of more than 500 million bot visits found only 408 requests for the file, which we covered in [do AI crawlers actually read llms.txt](/blog/do-ai-crawlers-read-llms-txt).

And reporting total bot hits as an AI visibility metric is worse than reporting nothing, because it moves in the wrong direction with confidence. Our own headline grew 884% across the seven weeks while the citing traffic fell 70%.

For context on what happens after retrieval, our concluded 63-day CITE Index study of 90,132 AI answers found ChatGPT cited a source in 92.5% of its answers and Google AI Mode in 97.4%. The engines are citing. The question is whether they ever came to read you. Full numbers are in our [AI search statistics](/ai-search-statistics).

## FAQ

### What are AI crawlers?

AI crawlers are automated bots that fetch web pages for AI systems. They fall into three jobs: training crawlers that collect content for model training, index crawlers that build the candidate pool for AI search, and answer-time fetchers that read a page the moment a user asks a question. Only the last two can produce a citation.

### Should I block AI crawlers?

Block training crawlers only if you have a rights or bandwidth reason. Blocking them will not improve or harm your citation rate. Never block answer-time or index bots such as ChatGPT-User, OAI-SearchBot, or PerplexityBot, because those are the only ones that can put your brand in an answer.

### Which AI crawler sends the most traffic?

On our site, Meta-ExternalAgent sent 82.2% of all AI bot traffic in the week ending July 27, 2026, and it is a pure training crawler with no citation path. Volume leadership and citation value are close to unrelated. Report them as separate numbers.

### How do I see AI crawlers in my logs?

Match the user-agent string at the edge or in server logs, then write each hit to a persistent store. Most hosting platforms retain runtime logs for under an hour, so a weekly report built on live tailing will read zero. Our [AI crawler log audit guide](/blog/ai-crawler-log-audit-retrieval) covers the full workflow.

### Do AI crawlers send referral traffic?

Training crawlers send none by design. Answer-time fetches can produce a referral if the engine links the citation and the user clicks. Published crawl-to-referral ratios are lopsided: roughly 23,951 pages crawled per referral for ClaudeBot against 4.9 for traditional Google search, per [aggregated Cloudflare figures](https://www.digitalapplied.com/blog/ai-crawler-bot-traffic-statistics-2026-data-reference).

## Where to start this week

Pull one week of logs and classify every AI user agent into the three buckets. Do not clean it up, do not annotate it, just get the three totals.

If your answer-time number is under 5% of the total, which it probably is, you now know that almost everything you were calling AI crawler traffic was never going to cite you. That is a better place to start than a growth chart.

### Find out which AI crawlers are actually reading you at answer time

Cite Solutions instruments your logs, separates the bots that can cite you from the ones burning your bandwidth, and runs the content work that gets the citing bots to quote you.

[Book a Discovery Call](/contact)

Tags

[GEO](/tag/geo)[AEO](/tag/aeo)[AI citations](/tag/ai-citations)[AI visibility](/tag/ai-visibility)[ai search optimization](/tag/ai-search-optimization)[technical SEO](/tag/technical-seo)[AI retrieval](/tag/ai-retrieval)[how to](/tag/how-to)

## Continue the brief

[01Technical GuidesDoes Technical SEO Still Matter for AI Search?Technical SEO decides whether AI can read your site at all. Most AI crawlers skip JavaScript, so client-rendered pages stay invisible. Here is the fix.Jul 16, 2026Read→](/blog/does-technical-seo-matter-for-ai-search)[02Technical GuidesDo AI Crawlers Actually Read llms.txt?Across 500M+ AI bot visits, only 408 fetched llms.txt. Here is what AI crawlers really do with the file, and whether it is worth publishing.Jun 7, 2026Read→](/blog/do-ai-crawlers-read-llms-txt)[03StrategyHow to Optimize for Agentic Search in 2026Agentic search does not read one page. It fans out, opens sources, and checks claims. Here is how to optimize for agentic search in 2026.Jul 26, 2026Read→](/blog/agentic-search-how-to-optimize)

[FrameworkLearn the CITE framework behind our GEO and AEO workSee how Comprehend, Influence, Track, and Evolve turn AI visibility into an operating system.](/framework)[ServicesExplore our managed GEO services and AEO execution modelAudit, prompt discovery, content execution, and ongoing monitoring tied to AI search outcomes.](/services)[AuditStart with an AI visibility audit before executionUnderstand prompt coverage, recommendation gaps, source mix, and where competitors are winning.](/ai-visibility-audit)

On this page

On this page

## Work with us on this

[LLM SEOGet cited by ChatGPT, Gemini, Claude, and Perplexity.Explore→](/llm-seo)[GEO AgencyManaged generative engine optimization for B2B brands.Explore→](/geo-agency)[AEO ServicesAnswer engine optimization: be the answer AI quotes.Explore→](/aeo-services)

## Ready to become the answer AI gives?

Book a 30-minute discovery call. We'll show you what AI says about your brand today. No pitch. Just data.

[Book a Discovery Call](/contact)
