Technical

GPTBot, ClaudeBot, PerplexityBot: Which AI Crawlers Should You Allow?

A guide to the main AI crawlers, the difference between training and search bots, and robots.txt rules that block training without hurting AI search visibility.

Many websites reacted to the rise of generative AI by blocking every AI-related bot in robots.txt. Some of those sites have since noticed something unwelcome: they rarely appear as sources in AI search answers. The reason is that "AI crawlers" are not one thing.

Three kinds of AI bots

  • Training crawlers collect content that may be used to train future models. Examples: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Meta-ExternalAgent, Bytespider. Google-Extended and Applebot-Extended are control tokens rather than separate crawlers: they tell Google and Apple whether content may be used for their AI models.
  • Search crawlers build indexes that AI search features retrieve from and cite. Examples: OAI-SearchBot (ChatGPT search), PerplexityBot, Claude-SearchBot, and the classic Googlebot and Bingbot, whose indexes power Google's AI features and Microsoft Copilot.
  • User-requested fetchers retrieve a page when a person asks an assistant about it. Examples: ChatGPT-User, Claude-User, Perplexity-User.

The common mistake

A rule like this blocks far more than training:

User-agent: *
Disallow: /

So does copying a long "block all AI bots" list without checking what each bot does. If OAI-SearchBot or PerplexityBot can't read your pages, those products have little to cite.

A balanced configuration

If you want to opt out of model training but stay eligible for AI search citations:

# Opt out of model training
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

# Everyone else, including AI search crawlers
User-agent: *
Allow: /

Sitemap: https://yourbrand.com/sitemap.xml

Whether to allow training is a business decision with no universal answer. Publishers with valuable proprietary content often block it; brands that want to be widely known often allow it. Just make the choice deliberately.

How robots.txt rules are evaluated

  • A crawler follows the group that names it specifically, and only falls back to User-agent: * if none does.
  • Within a group, the most specific (longest) matching path wins; if an Allow and a Disallow match equally, Allow wins.
  • No robots.txt (a 404) means everything is allowed.

Check yours in 10 seconds

Our free AI crawler access checker fetches your robots.txt and shows, for 17 AI user agents, whether they're allowed and which rule applies.

MA
Mohamed Aziz

Founder of Citewise AI. Writes about how AI assistants find, choose and cite brands.

See where AI recommends you, today.

Free workspace with 25 credits every month. No card required.