AI Crawlers
AI crawlers are the automated agents that AI companies use to fetch web pages, and they do so for three distinct purposes: collecting text that may be used to train models, building the search index that AI answers draw on, and retrieving a specific page live because a user asked a question that needs it. Each purpose has its own user agent, its own relationship to robots.txt, and a different consequence when a site blocks it, which is why "block AI bots" is no longer a single decision.
Three Purposes, Three Sets of Bots
The major operators now document separate bots per purpose. The table below reflects each operator's own crawler documentation as checked in October 2026.
| Operator | Training | Search index | User-initiated fetch |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Google-Extended (a robots.txt control token with no user agent string of its own; covers Gemini training and grounding) | Googlebot (the Search index, which AI Overviews and AI Mode are built on) | Not covered here | |
| Perplexity | None documented (PerplexityBot is "not used to crawl content for AI foundation models") | PerplexityBot | Perplexity-User |
The volumes are lopsided. Cloudflare's Radar 2025 Year in Review reported that training crawling ran at seven to eight times the volume of search crawling at its peak, while user-action crawling, the smallest category, grew by more than 15x during 2025.
Which Bots Honour robots.txt
Training and search-index bots are the ones operators commit to controlling through robots.txt. OpenAI states that disallowing GPTBot signals a site's content should not be used for training, and recommends allowing OAI-SearchBot for sites that want to appear in ChatGPT search. Anthropic says its bots honour "industry standard directives in robots.txt" and also support Crawl-delay. Google-Extended is honoured as a token rather than a crawler, and Google's documentation says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."
User-initiated fetchers are the exception. OpenAI's wording for ChatGPT-User is that "because these actions are initiated by a user, robots.txt rules may not apply," and Perplexity says Perplexity-User "generally ignores robots.txt rules" for the same reason. Anthropic is the outlier among the three: it documents a robots.txt opt-out for Claude-User, while warning that using it "may reduce your site's visibility." A site that needs a hard block on user-triggered fetches therefore cannot rely on robots.txt alone.
Default Blocking at the Network Layer
Control is also moving from the site to the network in front of it. Cloudflare announced on July 1, 2026 that for new domains onboarding from September 15, 2026, bots it classifies as Training and Agent are blocked by default on pages that display ads, while Search bots remain allowed by default. The same post notes that multi-purpose crawlers are allowed or blocked according to all of their behaviours, so a customer who chooses to block Training also affects crawlers that combine training with search. Site owners on a CDN should check what the network is doing before assuming their robots.txt is the whole policy.
Blocking and AI Visibility
The evidence that blocking costs visibility is real but observational. A SIGIR 2026 study of 11,500 queries (Grossman et al.) found that "websites that block Google's AI crawler are significantly less likely to be retrieved by AIOs", meaning AI Overviews. That sits awkwardly beside Google's statement that Google-Extended has no effect on Search, and the paper does not establish cause: sites that block AI crawlers may differ in other ways. The operators' own documents point the same direction for search and user-fetch bots, since OpenAI ties OAI-SearchBot access to appearing in results and Anthropic ties Claude-User access to visibility. The defensible reading is that blocking training bots is a licensing choice with no documented visibility penalty from the operators, while blocking search-index or user-fetch bots removes a site from the pool those answers are drawn from.
JavaScript and Practice
Operator documentation is largely silent on rendering; OpenAI's bot page, for instance, does not mention JavaScript at all. The best public evidence is a Vercel analysis of its own network traffic from December 2024, which concluded that "none of the major AI crawlers currently render JavaScript", naming OpenAI's, Anthropic's and Perplexity's bots, and observed that ChatGPT and Claude crawlers fetch JavaScript files without executing them. The same analysis noted that Gemini relies on Googlebot's infrastructure, which does render. That finding is nearly two years old and has not been confirmed by the operators, so it should be treated as a strong default assumption rather than a guarantee: content that only exists after client-side rendering may be invisible to most AI crawlers.
Managing all of this by hand means maintaining a growing list of user agents. It is increasingly a platform feature instead: LightCMS, the CMS serving this site, exposes separate switches for training, AI search and user-fetch crawlers, generates the matching robots.txt, and reports which AI crawlers actually read the site.
Further Reading
- Overview of OpenAI Crawlers — OpenAI, accessed October 2026
- Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic, accessed October 2026
- Google's common crawlers (Googlebot, Google-Extended) — Google Search Central, accessed October 2026
- Perplexity Crawlers — Perplexity, accessed October 2026
- Content Independence Day: new AI bot defaults — Cloudflare, July 2026
- Cloudflare Radar 2025 Year in Review — Cloudflare, December 2025
- How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews — Grossman et al., SIGIR 2026
- The rise of the AI crawler — Vercel, December 2024