How to block or allow AI crawlers in robots.txt

Decide separately about two kinds of AI crawlers: those that fetch pages to cite them in AI answers, and those that collect text to train models. Blocking the first kind removes your site from ChatGPT search, Perplexity and similar answers. Blocking the second kind keeps your content out of future training sets and does not affect search or answers.

Many sites block “all AI” with one list and lose both. Check what yours does with the AI crawler checker.

The crawlers that matter

User-agent Company Purpose
OAI-SearchBot OpenAI Finds and cites pages in ChatGPT search
ChatGPT-User OpenAI Fetches a page when a ChatGPT user asks it to
GPTBot OpenAI Collects training data
Claude-SearchBot Anthropic Search for Claude's answers
Claude-User Anthropic Fetches a page when a Claude user asks it to
ClaudeBot Anthropic Collects training data
PerplexityBot Perplexity Indexes pages for Perplexity answers
Perplexity-User Perplexity Fetches a page for a user's question
Google-Extended Google Not a crawler: a token that controls use in Gemini models. Does not affect Google Search
Googlebot Google Google Search, including AI Overviews
Bingbot Microsoft Bing search, which also feeds Copilot
Applebot-Extended Apple Controls use of content for Apple's AI models
CCBot Common Crawl Open web archive widely used for training

Vendors add and rename crawlers; check their documentation before relying on a name.

Rules you can copy

Stay in AI answers, keep out of training:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

Keep a section out of AI tools but in Google:

User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Disallow: /members/

Allow everything: no rules for these user agents, or no robots.txt at all.

How robots.txt matching works

  • A crawler follows only the most specific group that names it. If you add a User-agent: GPTBot group, GPTBot ignores your User-agent: * rules entirely, so repeat anything from * that should still apply.
  • Within a group, the longest matching path wins; Allow wins a tie.
  • Several User-agent lines can share one set of rules, as in the second example.

What robots.txt cannot do

  • It is a request. Reputable crawlers follow it; others may not. To enforce a block, use your firewall or CDN's bot controls.
  • It does not remove content that was already collected.
  • It does not stop a page from being shown if someone pastes its text into a chatbot.

Do not block Googlebot to stop AI Overviews

Google's AI Overviews use the normal search index. Blocking Googlebot removes the site from Google Search entirely; blocking Google-Extended does not affect AI Overviews. Controls such as nosnippet limit what Google can show from a page, in AI Overviews and in normal results alike.

What SignalCrawler checks here

The audit reads robots.txt as each crawler would and reports when AI search and answer crawlers are blocked, separately from training crawlers, with the line that blocks them. Blocking AI search crawlers is labelled as a decision for the site owner, not an error, because it can be deliberate.