sitemap.xml for AI-search crawlers
Whether GPTBot, PerplexityBot, ClaudeBot, and Google-Extended
actually read sitemap.xml — and the 6 things you
need to do so they discover your best pages instead of guessing.
Do AI crawlers actually read sitemaps?
Short answer: yes. Here's what each of the four AI crawlers that matter today does when it hits your domain:
| Crawler | Reads sitemap.xml? | Behaviour |
|---|---|---|
GPTBot (OpenAI) | Yes | Discovers via robots.txt Sitemap directive. Prioritises URLs with recent <lastmod>. |
PerplexityBot | Yes | Uses sitemap to decide what to crawl for its live retrieval index. Small crawl budget per domain — a clean sitemap gets more of your pages indexed. |
ClaudeBot (Anthropic) | Yes | Standard sitemap discovery. Respects <priority> as a weak signal on which URLs to fetch first. |
Google-Extended | Yes (via Googlebot) | Uses the same sitemap Google's classic crawler does. Gemini's live retrieval piggybacks on this. |
The catch: they use it as a discovery aid, not a ranking signal. Being in the sitemap doesn't get you cited — but being absent means AI crawlers may never find pages that aren't linked from your homepage. On sites over ~50 pages, this gap widens fast.
The copy-paste template
Put this at https://yourdomain.com/sitemap.xml.
Every field below is either required or high-signal —
<priority> and <changefreq>
are increasingly ignored, but <lastmod> is
the one field AI crawlers genuinely act on.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://yourdomain.com/</loc>
<lastmod>2026-08-27</lastmod>
</url>
<url>
<loc>https://yourdomain.com/pricing/</loc>
<lastmod>2026-08-20</lastmod>
</url>
<url>
<loc>https://yourdomain.com/blog/why-llms-txt/</loc>
<lastmod>2026-08-15</lastmod>
</url>
</urlset>
Then in your robots.txt:
User-agent: *
Allow: /
Sitemap: https://yourdomain.com/sitemap.xml Every field, explained
| Field | Required? | What AI crawlers use it for |
|---|---|---|
<loc> | Required | Absolute canonical URL, HTTPS, matches the served URL exactly (trailing slash, casing). Mismatches cause the crawler to fetch twice or skip entirely. |
<lastmod> | Critical | ISO 8601 date. AI crawlers with tight budgets skip URLs whose lastmod hasn't changed since their last visit. Wrong lastmod is worse than none. |
<changefreq> | Optional (mostly ignored) | Google deprecated this in 2023. AI crawlers followed. Skip it — it adds bytes for no benefit. |
<priority> | Optional | Weak signal. ClaudeBot uses it as a tiebreaker for crawl order; others ignore. Fine to include on your top ~10 pages at 0.9–1.0; don't bother otherwise. |
Sitemap indexes: when you have >50k URLs
Sitemap protocol caps each file at 50MB uncompressed and 50,000 URLs. Over that, split into multiple sitemaps and reference them from a sitemap index file.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://yourdomain.com/sitemap-pages.xml</loc>
<lastmod>2026-08-27</lastmod>
</sitemap>
<sitemap>
<loc>https://yourdomain.com/sitemap-blog.xml</loc>
<lastmod>2026-08-26</lastmod>
</sitemap>
</sitemapindex>
Point robots.txt at the index, not the
individual files. AI crawlers follow one level of indirection
reliably; two is fine but unnecessary.
Verify AI crawlers can actually read it
- Curl it as GPTBot.
curl -A "GPTBot" https://yourdomain.com/sitemap.xml— must return200,content-type: application/xml, and the full XML body. Anything else (403, HTML redirect, Cloudflare challenge page) means the crawler is being blocked. - Check robots.txt exposure.
curl https://yourdomain.com/robots.txt | grep -i sitemap— must show aSitemap:line pointing at the absolute HTTPS URL of your sitemap. - Validate the XML. Paste the URL at
xml-sitemaps.com/validate-xml-sitemap.html
— flags missing
<lastmod>, malformed dates, URLs that 404, and encoding issues. - Confirm lastmod accuracy. Pick 3 URLs from
the sitemap, fetch each, compare the
Last-ModifiedHTTP header (or the articledateModified) to the<lastmod>in the sitemap. Drift over a few days is fine; drift over months breaks crawler trust. - Run the free AICite audit. It flags missing sitemap.xml, unreachable sitemap URL, missing robots.txt declaration, and blocked-for-AI-crawlers responses as part of the 13-signal AI-search grade.
The 5 mistakes that break AI-crawler discovery
- Serving sitemap.xml over HTTP. Modern AI
crawlers hard-fail on non-HTTPS. Redirect
http://yourdomain.com/sitemap.xmlto HTTPS with a 301. - Content-type
text/html. Cloudflare Pages and Vercel occasionally serve XML with the wrong content-type when a rewrite catches the request. Fix by adding an explicit route header:content-type: application/xml; charset=utf-8. - Broken
<lastmod>. Common bugs: using ISO 8601 with time and no timezone (2026-08-27T14:30:00— reject), using US date format (08/27/2026— reject), or setting every URL'slastmodto build time so nothing ever looks stable. UseYYYY-MM-DDmatching the article's actual last edit. - Including dead or noindex URLs. Every 404
in your sitemap costs crawler trust. Every URL with
<meta name="robots" content="noindex">is a wasted fetch. Filter aggressively. - No
Sitemap:in robots.txt. Some crawlers guess/sitemap.xml. Some don't. The one-line declaration is free — always ship it.
Check your site now
FAQ
Do I still need sitemap.xml if I have llms.txt?
Yes. sitemap.xml is machine-readable inventory
for crawlers — it tells them what URLs exist and when they
changed. llms.txt is a curated human-readable
index for LLMs at inference time — it tells them what your
best content is. Different consumers, different purposes.
Ship both; they don't overlap.
Should I gzip sitemap.xml?
Optional. All major AI crawlers accept .xml.gz
and follow the Content-Encoding: gzip HTTP
header. Worth doing over ~5MB uncompressed. Under that,
skip the extra build step.
Does sitemap.xml help with Perplexity citations?
Indirectly. PerplexityBot has a small per-domain crawl budget
and heavily prefers URLs it can discover cheaply. A clean
sitemap with accurate lastmod lets Perplexity
find and re-fetch your best pages inside its budget. Without
one, deep pages (blog posts more than 2 clicks from your
homepage) often go uncrawled.
Can I block AI crawlers from sitemap.xml specifically?
Not cleanly. robots.txt can block a crawler
from fetching the sitemap, but that only stops discovery
for well-behaved crawlers; it doesn't hide the URL list
from anyone who guesses /sitemap.xml. If you
want to hide URLs from AI training, use robots.txt to block
the crawler from your content pages, not from the
sitemap file.
What about video sitemaps and image sitemaps?
AI crawlers currently ignore the extended sitemap namespaces for video, image, and news. Google's classic search still uses them, so keep them if you already have them. Not worth adding solely for AI-search readiness.
Ship sitemap.xml once and regenerate it on every
deploy. Combined with llms.txt (curated index)
and a clean robots.txt (crawler policy), it's the
third of the three files every AI-search-ready domain serves
from its root. For the full 13-signal fix pack, see the
AICite Pro Report ($24, 30 seconds).