Enter a site and see, for each major crawler, whether its robots.txt allows the homepage, and which line decides it. The checker reads robots.txt with the same matching rules Google uses: the most specific user-agent group applies, and within it the longest matching rule wins.
Two kinds of AI crawlers
The difference matters more than any single bot name.
Search and answer crawlers fetch pages to quote and link them in AI answers: OAI-SearchBot and ChatGPT-User for ChatGPT, PerplexityBot and Perplexity-User, Claude-SearchBot and Claude-User, Applebot for Siri and Spotlight, Bingbot for Bing and Copilot. Blocking them removes your site from those answers and from the visits they send.
Training crawlers collect text to train models: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (a token that controls use in Gemini, not a separate crawler), Applebot-Extended, CCBot (Common Crawl), Bytespider, Meta-ExternalAgent and Amazonbot. Blocking them does not remove you from search or AI answers.
Googlebot crawls for Google Search, and Google's AI Overviews use the same index. Blocking Googlebot removes you from Google entirely; blocking Google-Extended does not affect Search.
Reading the result
- Allowed, no rule applies: the crawler is not mentioned and the
*group does not block it. - Allowed by a rule: an
Allowline in the crawler's own group or in*. - Blocked: the
Disallowline that matches the homepage is shown with its line number. - No robots.txt: every crawler may crawl everything. That is fine if it is what you want.
The checker tests the homepage. A rule can still block a section such as /blog/; read the file to be sure.
Changing what you allow
To block training while staying in AI answers:
User-agent: GPTBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Leave OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot without such rules. The full guide, with the reasoning and caveats, is how to block or allow AI crawlers in robots.txt.
robots.txt is a request, not a lock: reputable crawlers follow it, and others may not. It also does not remove content that was already collected.