What is llms.txt? The AI-era sitemap standard, explained
llms.txt is a plain-text file at the root of your site that
tells ChatGPT, Perplexity, Claude, and Google AI Overviews which of your
pages matter and how they're structured. Think of it as
sitemap.xml for the LLM era — same idea, different reader.
What llms.txt actually is
llms.txt was proposed in
September 2024 by Jeremy Howard
as a machine-readable index of the pages on a site that are worth
ingesting into a large language model. It solves a specific problem:
LLMs have small context windows and web content is optimised for humans
(nav, ads, tracking, hero images). Without a curated index, an LLM
either crawls too much irrelevant HTML or misses the important bits.
The format is deliberately simple — it's markdown. Here's a minimal example:
# AICite
> Free A–F audit for how well your site is set up to be cited by
> ChatGPT, Perplexity, Claude, and Google AI Overviews.
## Core pages
- [Home](https://aicite.dev/): Free audit tool and overview
- [Pro Report](https://aicite.dev/report/): $24 fix pack with copy-paste snippets
- [What is llms.txt?](https://aicite.dev/why-llms-txt/): This article
## Reference
- [llmstxt.org spec](https://llmstxt.org/)
Three parts: a # H1 with your site name, an optional
> blockquote summary, and one or more
## H2 sections listing pages as markdown links with
a colon-and-description after each link. That's it.
Why it matters right now
The share of clicks going to LLM answer engines is climbing every quarter. Every time someone asks ChatGPT or Perplexity a question adjacent to your product, one of two things happens:
- The model has ingested your site and cites you.
- The model has not, and cites your competitor.
llms.txt is the cheapest, highest-leverage way to push the
needle toward option 1. Unlike SEO — which requires backlinks, page
rank, and months of patience — llms.txt is a static file
you ship once and update when your content changes. There is no
algorithm to game and no ranking penalty to worry about.
llms.txt vs sitemap.xml vs robots.txt
Three files, three jobs. They don't overlap; you want all three:
| File | Reader | Purpose |
|---|---|---|
robots.txt | Any crawler | Tells crawlers what they may and may not fetch. AI-specific user-agents (GPTBot, ClaudeBot, PerplexityBot) go here. |
sitemap.xml | Search engine crawlers | Lists every URL the search engine should index, with last-modified dates. |
llms.txt | LLM answer engines | Curated shortlist of the pages worth ingesting, in markdown, with human-written descriptions. |
How to ship one in 10 minutes
Want a head start? The free llms.txt generator drafts a valid starter file from your homepage — then just edit the links below.
- Pick 5–20 key pages. Home, pricing, docs landing, top blog posts, key product pages. Skip the newsletter archive.
- Write a one-line description of each. Not a keyword-stuffed meta description — a plain sentence a human wrote.
- Group them into 2–4 sections. "Core pages", "Reference", "Guides", "About" — whatever fits your site.
- Save the result as
llms.txtin your public/static folder so it's served at/llms.txt. - Verify it works by running an audit — the free AICite check looks for it in ~2 seconds and tells you exactly what's wrong if the file isn't reachable.
Optional: llms-full.txt
For docs sites and knowledge bases, ship a companion
llms-full.txt that contains the full markdown-converted
body of each page listed in llms.txt, concatenated
together. This lets an LLM ingest your whole site in a single request.
It's larger — often multiple MB — but it makes your content trivial
for RAG pipelines to pick up.
Common mistakes
- Serving it as HTML. The file must be
text/plainortext/markdown. Some frameworks route unknown paths to an HTML 404 — check withcurl -I. - Only shipping it on the apex domain. There is no
subdomain fallback. If you have
docs.example.comit needs its own file. - Linking to gated pages. An LLM crawler is not going to log in. Link only to pages a public HTTP request can fetch.
- Copy-pasting your sitemap. A sitemap has thousands
of URLs.
llms.txtshould have the 5–50 pages that actually matter, with human-written descriptions.
Check whether your site is AI-ready
FAQ
Do I need llms.txt if I already have great SEO?
Yes. Search-engine ranking signals (backlinks, page rank, content quality) are different from LLM-ingestion signals (curated index, structured markdown, machine-readable schema). Ranking #1 in Google doesn't guarantee ChatGPT cites you — the two systems read the web differently.
Will llms.txt hurt my SEO?
No. It's a separate file with a distinct name at the site root. Search engines ignore files they don't recognise, and the file is not linked from any HTML page (unless you choose to link it), so it has zero effect on page rank.
How often should I update it?
Whenever the shape of your site changes — new product page, new docs section, new pricing tier. Not on every blog post. Think of it as the equivalent of updating your top-nav.
Is the standard finalised?
Still "proposed" as of 2026, but adoption is broad enough (Anthropic, Perplexity, Hugging Face, many docs frameworks) that the shape you ship today is very unlikely to be wrong tomorrow.
Ship an llms.txt today. It's a 10-minute file that pays
rent for years as more traffic moves from search results to AI answers.
And if you'd like a full readiness check across 13 signals plus
ready-to-paste snippets, the
AICite Pro Report
is $24 and takes 30 seconds.