Skip to content
AICite

What is llms.txt? The AI-era sitemap standard, explained

llms.txt is a plain-text file at the root of your site that tells ChatGPT, Perplexity, Claude, and Google AI Overviews which of your pages matter and how they're structured. Think of it as sitemap.xml for the LLM era — same idea, different reader.

What llms.txt actually is

llms.txt was proposed in September 2024 by Jeremy Howard as a machine-readable index of the pages on a site that are worth ingesting into a large language model. It solves a specific problem: LLMs have small context windows and web content is optimised for humans (nav, ads, tracking, hero images). Without a curated index, an LLM either crawls too much irrelevant HTML or misses the important bits.

The format is deliberately simple — it's markdown. Here's a minimal example:

# AICite

> Free A–F audit for how well your site is set up to be cited by
> ChatGPT, Perplexity, Claude, and Google AI Overviews.

## Core pages

- [Home](https://aicite.dev/): Free audit tool and overview
- [Pro Report](https://aicite.dev/report/): $24 fix pack with copy-paste snippets
- [What is llms.txt?](https://aicite.dev/why-llms-txt/): This article

## Reference

- [llmstxt.org spec](https://llmstxt.org/)

Three parts: a # H1 with your site name, an optional > blockquote summary, and one or more ## H2 sections listing pages as markdown links with a colon-and-description after each link. That's it.

Why it matters right now

The share of clicks going to LLM answer engines is climbing every quarter. Every time someone asks ChatGPT or Perplexity a question adjacent to your product, one of two things happens:

  1. The model has ingested your site and cites you.
  2. The model has not, and cites your competitor.

llms.txt is the cheapest, highest-leverage way to push the needle toward option 1. Unlike SEO — which requires backlinks, page rank, and months of patience — llms.txt is a static file you ship once and update when your content changes. There is no algorithm to game and no ranking penalty to worry about.

llms.txt vs sitemap.xml vs robots.txt

Three files, three jobs. They don't overlap; you want all three:

File Reader Purpose
robots.txt Any crawler Tells crawlers what they may and may not fetch. AI-specific user-agents (GPTBot, ClaudeBot, PerplexityBot) go here.
sitemap.xml Search engine crawlers Lists every URL the search engine should index, with last-modified dates.
llms.txt LLM answer engines Curated shortlist of the pages worth ingesting, in markdown, with human-written descriptions.

How to ship one in 10 minutes

Want a head start? The free llms.txt generator drafts a valid starter file from your homepage — then just edit the links below.

  1. Pick 5–20 key pages. Home, pricing, docs landing, top blog posts, key product pages. Skip the newsletter archive.
  2. Write a one-line description of each. Not a keyword-stuffed meta description — a plain sentence a human wrote.
  3. Group them into 2–4 sections. "Core pages", "Reference", "Guides", "About" — whatever fits your site.
  4. Save the result as llms.txt in your public/static folder so it's served at /llms.txt.
  5. Verify it works by running an audit — the free AICite check looks for it in ~2 seconds and tells you exactly what's wrong if the file isn't reachable.

Optional: llms-full.txt

For docs sites and knowledge bases, ship a companion llms-full.txt that contains the full markdown-converted body of each page listed in llms.txt, concatenated together. This lets an LLM ingest your whole site in a single request. It's larger — often multiple MB — but it makes your content trivial for RAG pipelines to pick up.

Common mistakes

  • Serving it as HTML. The file must be text/plain or text/markdown. Some frameworks route unknown paths to an HTML 404 — check with curl -I.
  • Only shipping it on the apex domain. There is no subdomain fallback. If you have docs.example.com it needs its own file.
  • Linking to gated pages. An LLM crawler is not going to log in. Link only to pages a public HTTP request can fetch.
  • Copy-pasting your sitemap. A sitemap has thousands of URLs. llms.txt should have the 5–50 pages that actually matter, with human-written descriptions.

Check whether your site is AI-ready

FAQ

Do I need llms.txt if I already have great SEO?

Yes. Search-engine ranking signals (backlinks, page rank, content quality) are different from LLM-ingestion signals (curated index, structured markdown, machine-readable schema). Ranking #1 in Google doesn't guarantee ChatGPT cites you — the two systems read the web differently.

Will llms.txt hurt my SEO?

No. It's a separate file with a distinct name at the site root. Search engines ignore files they don't recognise, and the file is not linked from any HTML page (unless you choose to link it), so it has zero effect on page rank.

How often should I update it?

Whenever the shape of your site changes — new product page, new docs section, new pricing tier. Not on every blog post. Think of it as the equivalent of updating your top-nav.

Is the standard finalised?

Still "proposed" as of 2026, but adoption is broad enough (Anthropic, Perplexity, Hugging Face, many docs frameworks) that the shape you ship today is very unlikely to be wrong tomorrow.

Ship an llms.txt today. It's a 10-minute file that pays rent for years as more traffic moves from search results to AI answers. And if you'd like a full readiness check across 13 signals plus ready-to-paste snippets, the AICite Pro Report is $24 and takes 30 seconds.

Related: robots.txt for AI crawlers — GPTBot, ClaudeBot, PerplexityBot · JSON-LD Organization schema for AI-search · State of AI-search 2026: 25 sites ranked · sitemap.xml for AI crawlers · OpenGraph for AI-search · llms.txt vs robots.txt vs sitemap.xml.