robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot
The 2026 field guide to configuring robots.txt for AI
answer engines and training crawlers. Copy-paste snippets to
allow all, block all, or mix — and how to verify it's actually
working.
The AI-crawler user-agents that matter in 2026
These are the five you almost certainly want to name in
robots.txt. All of them publicly commit to honouring
the Robots Exclusion Standard.
| User-agent | Operator | Powers |
|---|---|---|
GPTBot | OpenAI | ChatGPT browsing + model training |
ClaudeBot | Anthropic | Claude answering + model training |
PerplexityBot | Perplexity | Perplexity answer engine |
Google-Extended | Gemini + AI-answer training (not search) | |
CCBot | Common Crawl | Public dataset that feeds many open LLMs |
Note that Googlebot is not in this list.
Blocking Googlebot hurts your normal search ranking.
Google-Extended is a separate agent Google introduced
in September 2023 specifically so publishers could opt out of AI
training without opting out of search.
Snippet: allow all AI crawlers
Best default if you want to appear in AI answers. Copy this into
the bottom of your existing robots.txt:
# AI answer engines & training crawlers — explicitly allowed
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: CCBot
Allow: /
The Allow: / is technically redundant (the default is
allow) but it makes your intent unambiguous to any tool auditing your
site — including ours.
Snippet: block all AI crawlers
For paywalled publishers or anyone treating their content as product. This blocks answer engines and training crawlers without touching search:
# AI answer engines & training crawlers — explicitly blocked
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: / Snippet: allow answer engines, block training
A pragmatic middle position: let AI answer engines cite your content in real time (so you get referral traffic), but block crawlers whose primary purpose is bulk training-data collection:
# Allow real-time answer engines
User-agent: PerplexityBot
Allow: /
# ClaudeBot handles both — decide based on your priorities
User-agent: ClaudeBot
Allow: /
# Block training-oriented crawlers
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: / The line between "answering" and "training" is fuzzy — GPTBot and ClaudeBot power both — so pick the rule that matches your business goal, not the crawler's stated intent.
Snippet: block AI crawlers from specific folders only
For sites with mixed content — e.g. a marketing site you want indexed and a paywalled archive you don't:
User-agent: GPTBot
Disallow: /premium/
Disallow: /members/
Allow: /
User-agent: ClaudeBot
Disallow: /premium/
Disallow: /members/
Allow: /
User-agent: Google-Extended
Disallow: /premium/
Disallow: /members/
Allow: / How to verify it's working
- Fetch the file directly.
curl -A "GPTBot" -I https://yourdomain.com/robots.txt
Should return200 OKwith content-typetext/plain. - Confirm the rules parse. Paste your file into TechnicalSEO's robots.txt tester — it accepts an arbitrary user-agent and tells you what would be allowed or blocked.
- Reference the sitemap. Add
Sitemap: https://yourdomain.com/sitemap.xmlat the top of the file. Both classical and AI crawlers respect this hint. - Run a free AICite audit. The free grade parses your
/robots.txtand tells you whether you have explicit AI-crawler rules or an implicit allow.
Common mistakes
- Blocking
Googlebotthinking you're blocking Gemini. You're not — you're blocking Google Search. UseGoogle-Extendedinstead. - Wildcard user-agent Disallow, expecting AI crawlers to honour it. They do, but you also just blocked Googlebot and everything else. Name AI crawlers explicitly.
- Serving robots.txt as HTML. It must be
text/plain. Framework 404 handlers sometimes return an HTML page from unknown paths — check withcurl -I. - Relying on robots.txt to hide sensitive content. robots.txt is a request, not a firewall. Anything you actually need private needs auth. The "robots.txt is not a security measure" rule applies double for AI crawlers.
- Forgetting the newline between blocks. Each
User-agentblock needs a blank line separator. Concatenated blocks are parsed as one big group and produce surprising results.
Check your site now
FAQ
Do I need one User-agent block per crawler, or can I list them together?
You can list multiple User-agent lines followed by a
single set of rules — they apply as a group. But naming each crawler
separately makes intent easier to audit and easier to differentiate
later. Prefer one block per crawler.
What about Bingbot and Amazonbot?
Bingbot powers Bing's Copilot answers as of 2024; if you want to appear there, don't block it. Amazonbot feeds Alexa answers. Both respect robots.txt. Include them if that's your audience.
How often should I update robots.txt?
Only when a rule genuinely changes — new user-agent, new folder policy. It's not versioned content. Well-behaved crawlers re-fetch it every ~24 hours.
Does the order of blocks matter?
For AI crawlers, no — they look up their own user-agent and use
the matching block. Convention is: Sitemap: at the
top, generic User-agent: * group, then specific
AI-crawler groups.
Configure the rules once and forget them — for most sites, the
maintenance cost of robots.txt is zero. What matters
is having an explicit rule so AI crawlers know what to do.
For the full 13-signal fix pack — llms.txt, JSON-LD,
OpenGraph, and the rest — see the
AICite Pro Report ($24, 30 seconds).