# llms.txt vs robots.txt

> llms.txt vs robots.txt: robots.txt tells crawlers what they may access; llms.txt offers agents a curated guide. Which bots honour each, and the usage data.

Source: https://metavert.io/compare/llms-txt-vs-robots-txt  
Published: 2026-10-07  
Updated: 2026-10-07

Comparison

**llms.txt** is a proposed Markdown file that gives language models and agents a curated guide to a site's most useful content, and **robots.txt** is the long-established file that tells automated crawlers which parts of a site they may access. [llms.txt](https://metavert.io/llms-txt) is an invitation; robots.txt is a set of permissions. They are not alternatives, and neither substitutes for the other.

The llms.txt proposal draws the line itself: “robots.txt lets automated tools know what access to a site is considered acceptable, such as for search indexing bots. llms.txt information is instead used on demand, when an agent needs information.” The practical difference in 2026 is adoption by the machines on the other end. The major [AI crawlers](https://metavert.io/ai-crawlers) document how they treat robots.txt, and getting it wrong can remove a site from [AI search](https://metavert.io/ai-search) results. No major AI search engine has said it uses llms.txt for retrieval or ranking, and log data shows most of these files are never requested. The file that is read is the one worth getting right first.

## Feature Comparison

| Dimension | llms.txt | robots.txt |
| --- | --- | --- |
| Purpose | Help agents find and use the right content | State which crawlers may access which paths |
| Status | Community proposal by Jeremy Howard, first published September 3, 2024 | IETF standard: RFC 9309, Robots Exclusion Protocol, September 2022 |
| Format | Markdown: an H1 title, optional summary, sections of annotated links | Plain-text user-agent groups with allow and disallow rules |
| Location | /llms.txt, optionally in subpaths such as /docs/llms.txt | /robots.txt at the root of the host |
| Controls access? | No | Expresses access rules; compliance is voluntary |
| Enforcement | None; purely informational | None in the protocol: “These rules are not a form of access authorization” |
| Honoured by AI index and training crawlers | No documented commitment from major engines | Yes: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot |
| User-initiated fetchers | May read it on demand | ChatGPT-User: rules “may not apply”; Perplexity-User “generally ignores” them |
| Google's position | No new “AI text files” needed for AI Overviews or AI Mode | Standard crawl control; Google-Extended token governs Gemini training and grounding |
| Observed usage | 97% of files received zero requests (Ahrefs, Jun 2026) | Fetched routinely by compliant crawlers as part of the protocol |
| Main consumers | Coding agents and agent infrastructure | Search, training and AI-search crawlers |
| Cost of a mistake | Low: a stale or ignored file | High: accidental removal from search or AI-search indexes |

## Detailed Analysis

### What Each File Controls

robots.txt speaks to crawlers before they fetch anything. It names user agents and lists paths they may or may not request. RFC 9309 is explicit about its limits: the rules “are not a form of access authorization,” and the protocol is no substitute for real content security. A well-behaved crawler obeys; a badly behaved one is not stopped by the file.

llms.txt controls nothing. It is a short Markdown document listing a site's key pages, often with links to clean [Markdown](https://metavert.io/markdown) versions, so that an agent with a limited context window can go straight to the useful material. It cannot grant or deny access, it cannot opt a site out of training, and a path omitted from llms.txt remains as crawlable as before. Using it to express restrictions is a category error.

### Which Bots Honour What

The robots.txt picture is reasonably well documented. OpenAI states that GPTBot (training) and OAI-SearchBot (search index) honour it. Anthropic's ClaudeBot, Claude-SearchBot and Claude-User all honour it, according to its crawler documentation as reported by Search Engine Roundtable in February 2026. Perplexity recommends allowing PerplexityBot in robots.txt to appear in its results. The exceptions are the fetchers acting on a person's request: OpenAI says of ChatGPT-User that “robots.txt rules may not apply,” and Perplexity says Perplexity-User “generally ignores robots.txt rules.”

The consequence is that robots.txt now carries three separate decisions per vendor: training, AI-search indexing, and user-initiated fetching. They can be answered differently. A site can block GPTBot and still allow OAI-SearchBot. Google handles the split with a token: Google-Extended governs Gemini training and grounding and, per Google, does not affect inclusion or ranking in Search. Platform defaults are shifting too. Cloudflare announced that from September 15, 2026, training and agent bots are blocked by default on ad-supported pages for new domains, while search bots stay allowed.

For llms.txt there is no comparable set of commitments. Google's guidance on AI features says site owners do not need “new machine readable files, AI text files, or markup” to appear in AI Overviews or AI Mode.

### The llms.txt Usage Data

Ahrefs' June 2026 study examined server logs for 137,210 domains from May 2026. It found that 97% of llms.txt files received zero requests. Among the requests that did occur, AI retrieval bots accounted for 1.1%, while coding agents and agent infrastructure accounted for 10.5%, the largest AI category. The study's summary: “Zero AI bots ‘go looking’ for llms.txt files that don't exist.”

The exception is developer documentation. Mintlify, a documentation host reporting on its own customers (July 2026; vendor data), said 66% of traffic to its hosted docs in July 2026 came from agents, up from 15.2% at the start of the year, and that adding llms.txt cut agent 404 errors by about 90% in its benchmark. That fits the Ahrefs breakdown. The file is useful where coding agents arrive needing to find an API reference quickly, a use case that overlaps with the [Model Context Protocol](https://metavert.io/model-context-protocol) and is compared in [llms.txt vs MCP](https://metavert.io/compare/llms-txt-vs-mcp). It has no demonstrated effect on consumer AI search visibility.

### In Practice

The risk profile is lopsided. An llms.txt file costs little and may be ignored. A robots.txt error can be expensive: a SIGIR 2026 study (Grossman et al.) found that sites blocking AI crawlers show reduced visibility in AI Overviews, and a blanket disallow aimed at training bots can catch search bots by accident. Bot traffic is also changing in composition. Cloudflare's 2025 Year in Review reported user-action crawling grew more than fifteenfold during 2025, and that training crawling ran at seven to eight times the volume of search crawling.

Maintaining both files by hand is error-prone, and it is increasingly handled by the publishing platform. [LightCMS](https://metavert.io/lightcms), the CMS serving this site, generates llms.txt and robots.txt from per-purpose settings for training, AI search and user-fetch crawlers, so the two stay consistent with one policy. Whatever the tooling, the policy itself has to be decided by a person: which uses of the content are acceptable, and for which companies.

## Best For

#### If you can only spend an hour on one this quarter

robots.txt

Audit it for rules that block AI-search crawlers unintentionally. This is the file compliant bots read, and an error here costs visibility.

#### If you want to opt out of AI training but stay in AI search

robots.txt

Disallow the training agents (GPTBot, ClaudeBot, Google-Extended) and allow the search agents (OAI-SearchBot, Claude-SearchBot, PerplexityBot). llms.txt cannot express any of this.

#### If you publish developer documentation or an API

llms.txt

Coding agents are the largest AI consumers of the file, and documentation hosts report fewer failed agent requests with it in place. This is the use case it was designed for.

#### If your goal is more citations in ChatGPT or AI Overviews

Neither

robots.txt can only make you eligible, and llms.txt has no demonstrated effect on consumer AI search. Citations depend on content and third-party mentions.

#### If you need to stop a bot that ignores the rules

Neither

Both files are advisory. Blocking requires enforcement at the server, firewall or CDN.

#### If your CMS generates both automatically

Both

Publish both and review the robots.txt output. The marginal cost of llms.txt is near zero and it serves agent visitors when they do arrive.

#### If you want to keep user-initiated fetchers out

robots.txt, with limits

Anthropic's Claude-User honours it. OpenAI and Perplexity say their user-triggered fetchers may not, so pair the rule with server-side controls.

## The Bottom Line

robots.txt and llms.txt answer different questions. One says what crawlers may access; the other suggests what agents should read. A site can sensibly publish both, and should not expect either to do the other's job.

They are not equally consequential. robots.txt is a ratified standard that the main AI crawlers say they follow, and it now encodes separate choices about training, AI-search indexing and user-initiated fetching. Errors in it have measurable effects on visibility. llms.txt is a proposal that, on current log evidence, almost nothing requests outside developer documentation, and that no major AI search engine has committed to using.

The sensible order of work is to get robots.txt right, enforce anything that matters at the server, and add llms.txt where agents are a real audience, which today mostly means documentation. Claims that llms.txt improves AI search rankings are not supported by the 2026 data.

## Related Topics

- [llms.txt](https://metavert.io/llms-txt)
- [AI Crawlers](https://metavert.io/ai-crawlers)
- [llms.txt vs MCP](https://metavert.io/compare/llms-txt-vs-mcp)
- [Model Context Protocol](https://metavert.io/model-context-protocol)
- [Agent Experience](https://metavert.io/agent-experience)
- [Markdown](https://metavert.io/markdown)
- [Generative Engine Optimization](https://metavert.io/generative-engine-optimization)
- [LightCMS](https://metavert.io/lightcms)
- [AI Search](https://metavert.io/ai-search)

## Further Reading

- [The /llms.txt File (proposal) – llmstxt.org](https://llmstxt.org/)
- [RFC 9309: Robots Exclusion Protocol (September 2022) – IETF](https://www.rfc-editor.org/rfc/rfc9309.html)
- [Overview of OpenAI Crawlers – OpenAI](https://developers.openai.com/api/docs/bots)
- [Anthropic Updates Its Crawler Docs (February 2026) – Search Engine Roundtable](https://www.seroundtable.com/anthropic-updates-its-crawler-docs-40978.html)
- [Perplexity Crawlers – Perplexity](https://docs.perplexity.ai/guides/bots)
- [AI Features and Your Website – Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)
- [Google Common Crawlers (Google-Extended) – Google Search Central](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)
- [llms.txt Study, 137,210 Domains (June 2026) – Ahrefs](https://ahrefs.com/blog/llmstxt-study/)
- [State of Docs Traffic (July 2026) – Mintlify](https://www.mintlify.com/blog/state-of-docs-traffic)
- [Content Independence Day: AI Options (July 2026) – Cloudflare](https://blog.cloudflare.com/content-independence-day-ai-options/)
- [Radar 2025 Year in Review – Cloudflare](https://blog.cloudflare.com/radar-2025-year-in-review/)
- [Grossman et al., Cross-Platform Source Overlap – SIGIR 2026 / arXiv](https://arxiv.org/abs/2604.27790)
