Skip to content
AICite

sitemap.xml for AI-search crawlers

Whether GPTBot, PerplexityBot, ClaudeBot, and Google-Extended actually read sitemap.xml — and the 6 things you need to do so they discover your best pages instead of guessing.

Do AI crawlers actually read sitemaps?

Short answer: yes. Here's what each of the four AI crawlers that matter today does when it hits your domain:

Crawler Reads sitemap.xml? Behaviour
GPTBot (OpenAI) Yes Discovers via robots.txt Sitemap directive. Prioritises URLs with recent <lastmod>.
PerplexityBot Yes Uses sitemap to decide what to crawl for its live retrieval index. Small crawl budget per domain — a clean sitemap gets more of your pages indexed.
ClaudeBot (Anthropic) Yes Standard sitemap discovery. Respects <priority> as a weak signal on which URLs to fetch first.
Google-Extended Yes (via Googlebot) Uses the same sitemap Google's classic crawler does. Gemini's live retrieval piggybacks on this.

The catch: they use it as a discovery aid, not a ranking signal. Being in the sitemap doesn't get you cited — but being absent means AI crawlers may never find pages that aren't linked from your homepage. On sites over ~50 pages, this gap widens fast.

The copy-paste template

Put this at https://yourdomain.com/sitemap.xml. Every field below is either required or high-signal — <priority> and <changefreq> are increasingly ignored, but <lastmod> is the one field AI crawlers genuinely act on.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://yourdomain.com/</loc>
    <lastmod>2026-08-27</lastmod>
  </url>
  <url>
    <loc>https://yourdomain.com/pricing/</loc>
    <lastmod>2026-08-20</lastmod>
  </url>
  <url>
    <loc>https://yourdomain.com/blog/why-llms-txt/</loc>
    <lastmod>2026-08-15</lastmod>
  </url>
</urlset>

Then in your robots.txt:

User-agent: *
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

Every field, explained

Field Required? What AI crawlers use it for
<loc> Required Absolute canonical URL, HTTPS, matches the served URL exactly (trailing slash, casing). Mismatches cause the crawler to fetch twice or skip entirely.
<lastmod> Critical ISO 8601 date. AI crawlers with tight budgets skip URLs whose lastmod hasn't changed since their last visit. Wrong lastmod is worse than none.
<changefreq> Optional (mostly ignored) Google deprecated this in 2023. AI crawlers followed. Skip it — it adds bytes for no benefit.
<priority> Optional Weak signal. ClaudeBot uses it as a tiebreaker for crawl order; others ignore. Fine to include on your top ~10 pages at 0.9–1.0; don't bother otherwise.

Sitemap indexes: when you have >50k URLs

Sitemap protocol caps each file at 50MB uncompressed and 50,000 URLs. Over that, split into multiple sitemaps and reference them from a sitemap index file.

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://yourdomain.com/sitemap-pages.xml</loc>
    <lastmod>2026-08-27</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://yourdomain.com/sitemap-blog.xml</loc>
    <lastmod>2026-08-26</lastmod>
  </sitemap>
</sitemapindex>

Point robots.txt at the index, not the individual files. AI crawlers follow one level of indirection reliably; two is fine but unnecessary.

Verify AI crawlers can actually read it

  1. Curl it as GPTBot. curl -A "GPTBot" https://yourdomain.com/sitemap.xml — must return 200, content-type: application/xml, and the full XML body. Anything else (403, HTML redirect, Cloudflare challenge page) means the crawler is being blocked.
  2. Check robots.txt exposure. curl https://yourdomain.com/robots.txt | grep -i sitemap — must show a Sitemap: line pointing at the absolute HTTPS URL of your sitemap.
  3. Validate the XML. Paste the URL at xml-sitemaps.com/validate-xml-sitemap.html — flags missing <lastmod>, malformed dates, URLs that 404, and encoding issues.
  4. Confirm lastmod accuracy. Pick 3 URLs from the sitemap, fetch each, compare the Last-Modified HTTP header (or the article dateModified) to the <lastmod> in the sitemap. Drift over a few days is fine; drift over months breaks crawler trust.
  5. Run the free AICite audit. It flags missing sitemap.xml, unreachable sitemap URL, missing robots.txt declaration, and blocked-for-AI-crawlers responses as part of the 13-signal AI-search grade.

The 5 mistakes that break AI-crawler discovery

  • Serving sitemap.xml over HTTP. Modern AI crawlers hard-fail on non-HTTPS. Redirect http://yourdomain.com/sitemap.xml to HTTPS with a 301.
  • Content-type text/html. Cloudflare Pages and Vercel occasionally serve XML with the wrong content-type when a rewrite catches the request. Fix by adding an explicit route header: content-type: application/xml; charset=utf-8.
  • Broken <lastmod>. Common bugs: using ISO 8601 with time and no timezone (2026-08-27T14:30:00 — reject), using US date format (08/27/2026 — reject), or setting every URL's lastmod to build time so nothing ever looks stable. Use YYYY-MM-DD matching the article's actual last edit.
  • Including dead or noindex URLs. Every 404 in your sitemap costs crawler trust. Every URL with <meta name="robots" content="noindex"> is a wasted fetch. Filter aggressively.
  • No Sitemap: in robots.txt. Some crawlers guess /sitemap.xml. Some don't. The one-line declaration is free — always ship it.

Check your site now

FAQ

Do I still need sitemap.xml if I have llms.txt?

Yes. sitemap.xml is machine-readable inventory for crawlers — it tells them what URLs exist and when they changed. llms.txt is a curated human-readable index for LLMs at inference time — it tells them what your best content is. Different consumers, different purposes. Ship both; they don't overlap.

Should I gzip sitemap.xml?

Optional. All major AI crawlers accept .xml.gz and follow the Content-Encoding: gzip HTTP header. Worth doing over ~5MB uncompressed. Under that, skip the extra build step.

Does sitemap.xml help with Perplexity citations?

Indirectly. PerplexityBot has a small per-domain crawl budget and heavily prefers URLs it can discover cheaply. A clean sitemap with accurate lastmod lets Perplexity find and re-fetch your best pages inside its budget. Without one, deep pages (blog posts more than 2 clicks from your homepage) often go uncrawled.

Can I block AI crawlers from sitemap.xml specifically?

Not cleanly. robots.txt can block a crawler from fetching the sitemap, but that only stops discovery for well-behaved crawlers; it doesn't hide the URL list from anyone who guesses /sitemap.xml. If you want to hide URLs from AI training, use robots.txt to block the crawler from your content pages, not from the sitemap file.

What about video sitemaps and image sitemaps?

AI crawlers currently ignore the extended sitemap namespaces for video, image, and news. Google's classic search still uses them, so keep them if you already have them. Not worth adding solely for AI-search readiness.

Ship sitemap.xml once and regenerate it on every deploy. Combined with llms.txt (curated index) and a clean robots.txt (crawler policy), it's the third of the three files every AI-search-ready domain serves from its root. For the full 13-signal fix pack, see the AICite Pro Report ($24, 30 seconds).

Related: What is llms.txt? · robots.txt for AI crawlers · JSON-LD Organization schema · State of AI-search 2026: 25 sites ranked · OpenGraph for AI-search · llms.txt vs robots.txt vs sitemap.xml.