Skip to content
AICite

robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot

The 2026 field guide to configuring robots.txt for AI answer engines and training crawlers. Copy-paste snippets to allow all, block all, or mix — and how to verify it's actually working.

The AI-crawler user-agents that matter in 2026

These are the five you almost certainly want to name in robots.txt. All of them publicly commit to honouring the Robots Exclusion Standard.

User-agent Operator Powers
GPTBot OpenAI ChatGPT browsing + model training
ClaudeBot Anthropic Claude answering + model training
PerplexityBot Perplexity Perplexity answer engine
Google-Extended Google Gemini + AI-answer training (not search)
CCBot Common Crawl Public dataset that feeds many open LLMs

Note that Googlebot is not in this list. Blocking Googlebot hurts your normal search ranking. Google-Extended is a separate agent Google introduced in September 2023 specifically so publishers could opt out of AI training without opting out of search.

Snippet: allow all AI crawlers

Best default if you want to appear in AI answers. Copy this into the bottom of your existing robots.txt:

# AI answer engines & training crawlers — explicitly allowed
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: CCBot
Allow: /

The Allow: / is technically redundant (the default is allow) but it makes your intent unambiguous to any tool auditing your site — including ours.

Snippet: block all AI crawlers

For paywalled publishers or anyone treating their content as product. This blocks answer engines and training crawlers without touching search:

# AI answer engines & training crawlers — explicitly blocked
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

Snippet: allow answer engines, block training

A pragmatic middle position: let AI answer engines cite your content in real time (so you get referral traffic), but block crawlers whose primary purpose is bulk training-data collection:

# Allow real-time answer engines
User-agent: PerplexityBot
Allow: /

# ClaudeBot handles both — decide based on your priorities
User-agent: ClaudeBot
Allow: /

# Block training-oriented crawlers
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

The line between "answering" and "training" is fuzzy — GPTBot and ClaudeBot power both — so pick the rule that matches your business goal, not the crawler's stated intent.

Snippet: block AI crawlers from specific folders only

For sites with mixed content — e.g. a marketing site you want indexed and a paywalled archive you don't:

User-agent: GPTBot
Disallow: /premium/
Disallow: /members/
Allow: /

User-agent: ClaudeBot
Disallow: /premium/
Disallow: /members/
Allow: /

User-agent: Google-Extended
Disallow: /premium/
Disallow: /members/
Allow: /

How to verify it's working

  1. Fetch the file directly.
    curl -A "GPTBot" -I https://yourdomain.com/robots.txt
    Should return 200 OK with content-type text/plain.
  2. Confirm the rules parse. Paste your file into TechnicalSEO's robots.txt tester — it accepts an arbitrary user-agent and tells you what would be allowed or blocked.
  3. Reference the sitemap. Add Sitemap: https://yourdomain.com/sitemap.xml at the top of the file. Both classical and AI crawlers respect this hint.
  4. Run a free AICite audit. The free grade parses your /robots.txt and tells you whether you have explicit AI-crawler rules or an implicit allow.

Common mistakes

  • Blocking Googlebot thinking you're blocking Gemini. You're not — you're blocking Google Search. Use Google-Extended instead.
  • Wildcard user-agent Disallow, expecting AI crawlers to honour it. They do, but you also just blocked Googlebot and everything else. Name AI crawlers explicitly.
  • Serving robots.txt as HTML. It must be text/plain. Framework 404 handlers sometimes return an HTML page from unknown paths — check with curl -I.
  • Relying on robots.txt to hide sensitive content. robots.txt is a request, not a firewall. Anything you actually need private needs auth. The "robots.txt is not a security measure" rule applies double for AI crawlers.
  • Forgetting the newline between blocks. Each User-agent block needs a blank line separator. Concatenated blocks are parsed as one big group and produce surprising results.

Check your site now

FAQ

Do I need one User-agent block per crawler, or can I list them together?

You can list multiple User-agent lines followed by a single set of rules — they apply as a group. But naming each crawler separately makes intent easier to audit and easier to differentiate later. Prefer one block per crawler.

What about Bingbot and Amazonbot?

Bingbot powers Bing's Copilot answers as of 2024; if you want to appear there, don't block it. Amazonbot feeds Alexa answers. Both respect robots.txt. Include them if that's your audience.

How often should I update robots.txt?

Only when a rule genuinely changes — new user-agent, new folder policy. It's not versioned content. Well-behaved crawlers re-fetch it every ~24 hours.

Does the order of blocks matter?

For AI crawlers, no — they look up their own user-agent and use the matching block. Convention is: Sitemap: at the top, generic User-agent: * group, then specific AI-crawler groups.

Configure the rules once and forget them — for most sites, the maintenance cost of robots.txt is zero. What matters is having an explicit rule so AI crawlers know what to do. For the full 13-signal fix pack — llms.txt, JSON-LD, OpenGraph, and the rest — see the AICite Pro Report ($24, 30 seconds).

Related: What is llms.txt? The AI-era sitemap standard, explained · JSON-LD Organization schema for AI-search · State of AI-search 2026: 25 sites ranked · sitemap.xml for AI crawlers · OpenGraph for AI-search · llms.txt vs robots.txt vs sitemap.xml.