llms.txt vs robots.txt vs sitemap.xml: which one do AI crawlers actually check?
Three files, three jobs. robots.txt is
permission. sitemap.xml is discovery.
llms.txt is curation. AI crawlers read them in
that order — and skipping any of them costs citations.
What each file actually does
| File | Role | Audience | Format |
|---|---|---|---|
robots.txt | Permission gate. Tells crawlers which paths they may fetch and which user-agents you welcome or block. | Every crawler — classical search, AI training, AI answer bots. | Plain text, RFC 9309. |
sitemap.xml | Discovery index. Enumerates every URL on the site so crawlers don't have to guess. | Search crawlers (Googlebot, Bingbot) and AI crawlers (GPTBot, ClaudeBot, PerplexityBot). | XML, sitemaps.org spec. |
llms.txt | Curated shortlist. A human-written list of the highest-value pages you want LLMs to read and cite. | LLM answer bots (Perplexity, ChatGPT browse, Claude tools) and training crawlers that respect the convention. | Markdown, llmstxt.org spec. |
The three files are complementary, not overlapping. They answer different questions: may I fetch?, what URLs exist?, and which pages matter most?
The order an AI crawler reads them
-
robots.txtfirst. Every compliant crawler fetcheshttps://yourdomain.com/robots.txtbefore doing anything else. If your rules disallow the crawler's user-agent (e.g.User-agent: GPTBot+Disallow: /), the crawler stops and reads nothing further. This is the permission gate. -
sitemap.xmlsecond. Once allowed, the crawler looks for aSitemap: https://yourdomain.com/sitemap.xmlline inrobots.txt, then falls back to the well-known location. It uses the sitemap to enumerate URLs without having to follow every link. -
llms.txtlast, as a quality signal. AI answer bots (Perplexity, some ChatGPT browsing sessions) fetch/llms.txtto see the shortlist. Pages in this file get pulled preferentially when the bot has to pick a limited set of URLs to include in its context window.
Skipping step 1 breaks compliance (crawlers may still fetch, but you have no control). Skipping step 2 slows discovery and can cost you inclusion in Bing and Google AI Overviews. Skipping step 3 costs you AI-answer citations that would otherwise land on your best pages instead of random archive-tagged posts.
Which AI crawlers read which file?
| Crawler | robots.txt | sitemap.xml | llms.txt |
|---|---|---|---|
| GPTBot (OpenAI training) | Yes — respects Disallow | Yes — follows Sitemap: line | Emerging — observed fetches in 2026 |
| ClaudeBot (Anthropic training) | Yes — respects Disallow | Yes | Yes — treats as curated input |
| PerplexityBot (Perplexity answers) | Yes | Yes | Yes — high signal, feeds citation ranking |
| Google-Extended (AI Overviews) | Yes | Yes — via Googlebot sitemap fetch | Not officially, but tracked internally |
| ChatGPT-User (browsing on user request) | Partial — read but not always enforced | Rarely fetched (URL-driven) | Yes — pulled when the domain is cited |
| Applebot-Extended (Apple Intelligence) | Yes | Yes | Emerging |
| CCBot (Common Crawl → training data) | Yes | Yes | Not yet |
Bottom line: robots.txt is universal, sitemap.xml is near-universal, and llms.txt is the growing frontier. Every major AI crawler either reads llms.txt today or is on a trajectory to.
Decision tree: which file fixes my problem?
| Problem | Fix in |
|---|---|
| An AI crawler is hammering my site and I want it to slow down or stop. | robots.txt — add Crawl-delay or Disallow: / for its user-agent. |
| My content is being used to train models and I want to opt out. | robots.txt — disallow GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended. |
| New pages take weeks to appear in search / AI answers. | sitemap.xml — add all URLs with <lastmod>, reference from robots.txt. |
| My homepage gets cited but my best deep pages never do. | llms.txt — list your deep pages with one-line descriptions so LLMs pull them first. |
| Perplexity cites a stale archive page instead of my latest post. | llms.txt — feature the latest post in the primary section. |
| Bingbot / Googlebot can't find my paginated blog archive. | sitemap.xml — split into a sitemap index if you have >50k URLs. |
| An unknown AI bot is fetching pages I don't want in training. | robots.txt — start with User-agent: * + Disallow: /private/, then add per-bot rules once you identify them. |
Common misconceptions
- "llms.txt replaces robots.txt for AI crawlers." False. llms.txt has no permission semantics. If you want to block a crawler, you must use robots.txt. llms.txt is a recommendation layer on top of permitted access.
- "If I disallow GPTBot in robots.txt, my site is safe from training." Partially true. Reputable bots (GPTBot, ClaudeBot, CCBot, Google-Extended) respect Disallow. Non-compliant scrapers ignore it. robots.txt is a contract, not a firewall.
- "sitemap.xml is only for Google." False. Every major AI crawler uses sitemap.xml to enumerate URLs. Skipping it means new pages take longer to appear in AI answers.
- "llms.txt only helps if my content is in the training set." False. Answer bots (Perplexity, ChatGPT browsing) fetch llms.txt at query time, not just at training time. It influences which pages get cited today, not just which pages inform future model updates.
- "You can't have both llms.txt and sitemap.xml." You should have both. Sitemap is for machines enumerating every URL; llms.txt is a curated shortlist for LLMs. They coexist without conflict.
Minimal versions of each file
robots.txt
# Allow everything by default
User-agent: *
Allow: /
# AI training bots — allow (change to Disallow: / to opt out)
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Google-Extended
Allow: /
Sitemap: https://yourdomain.com/sitemap.xml sitemap.xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://yourdomain.com/</loc>
<lastmod>2026-08-31</lastmod>
</url>
<url>
<loc>https://yourdomain.com/pricing/</loc>
<lastmod>2026-08-15</lastmod>
</url>
</urlset> llms.txt
# Acme Corp
> Real-time inventory tracking API for small retailers.
> Docs, pricing, and integration examples.
## Primary content
- [Getting started](https://acme.example/docs/getting-started/): 5-minute quickstart with a working example.
- [API reference](https://acme.example/docs/api/): complete endpoint list with request/response schemas.
- [Pricing](https://acme.example/pricing/): plans, limits, and enterprise options.
## About
- [About Acme](https://acme.example/about/): team, story, and mission. Each file has a dedicated deep-dive on this site if you want the full field guide: llms.txt, robots.txt, sitemap.xml.
Verify all three are being fetched
- Fetch each file as a real AI crawler.
curl -A "PerplexityBot" https://yourdomain.com/robots.txt, then swap in/sitemap.xmland/llms.txt. All three must return200. Any404or403silently costs citations. - Confirm the sitemap is referenced from robots.txt.
grep -i sitemap robots.txt— theSitemap:line is how most crawlers discover the sitemap without probing every well-known path. - Check server logs for AI-crawler user-agents.
Look for
GPTBot,ClaudeBot,PerplexityBot,Google-Extended. Frequency > 0 for all three files is the pass signal. - Run the AICite audit. It flags missing files, missing Sitemap directive, and llms.txt format issues in one A–F grade.
Check your site now
FAQ
Do I need to submit llms.txt anywhere, like I submit sitemap.xml to Google Search Console?
No. There is no llms.txt submission portal. AI crawlers
fetch /llms.txt at the well-known location
automatically the next time they visit your domain. Ship
it and it takes effect within days.
What if my sitemap.xml is huge — should I list everything in llms.txt too?
No. llms.txt is a shortlist, not a mirror of sitemap.xml. Aim for 10–50 URLs — the pages you'd hand-pick if a journalist asked "what should I read to understand your company?" Everything else stays in the sitemap.
Which file should I update most often?
sitemap.xml — it should update every time you
publish a page (most static site generators handle this
automatically). llms.txt updates when your
flagship content changes, maybe monthly.
robots.txt updates rarely, only when
permissions change.
Does llms.txt help with classical SEO (Google search results)?
Indirectly. Googlebot doesn't officially use llms.txt for classical ranking. But Google-Extended (the AI Overviews crawler) does track it, and Overview citations increasingly drive traffic that used to come from classical result position 1. Investing in llms.txt is investing in the category of search that's growing.
What happens if my three files contradict each other?
robots.txt always wins on permission — if it
says Disallow: /docs/, the crawler will not
fetch /docs/ even if that URL appears in
sitemap.xml or llms.txt. So don't
list disallowed URLs in the discovery or curation files —
it wastes crawl budget and looks broken in tooling.
Three files, three roles: robots.txt for
permission, sitemap.xml for discovery,
llms.txt for curation. Ship all three correctly
and you've covered the foundation of AI-search readiness.
For the full 13-signal fix pack, see the
AICite Pro Report ($24, 30 seconds).