Skip to content
ZeroServer.tools

AI Bot Access Tester

Test which AI crawlers your robots.txt allows or blocks. Supports all major bots and RFC 9309 wildcards.

Status: ALLOWEDMatched Agent: NoneMatched Rule: Default Allow

How robots.txt AI bot control works

Websites use robots.txt to tell crawlers which pages they may or may not access. AI training crawlers like GPTBot, ClaudeBot, and Google-Extended respect these rules. A Disallow: /under a bot's User-agent blocks all access; an empty Disallow or Allow: / permits it. The most specific matching rule wins per RFC 9309.

Built and maintained by Meet Shah · Last updated

What this tool is used for

  • Checking which AI crawlers your current rules allow or block.
  • Testing a new rule before deploying it, since these user-agents change often.
  • Confirming that blocking one AI crawler has not accidentally blocked a search crawler.
  • Auditing an inherited robots.txt for rules nobody documented.
  • Deciding on a policy by seeing what the current file actually says.

Frequently Asked Questions

How is this different from testing search crawlers?
Different user-agents with different purposes. GPTBot, ClaudeBot, Google-Extended, PerplexityBot and CCBot are about TRAINING and retrieval, not search ranking — so blocking them does not affect your search visibility.
What does Google-Extended actually do?
It is a control token, not a crawler. Blocking it opts your content out of Gemini and AI training WITHOUT affecting Googlebot or your search ranking — which is exactly why it exists as a separate token.
Does blocking these remove content already used?
No. robots.txt is forward-looking only, so anything already crawled and used in training is unaffected. Blocking prevents future collection and nothing more.
Will blocking AI bots hurt referral traffic?
Possibly, and it is a genuine trade-off. Some crawlers serve live retrieval that cites and links back — blocking them removes your content from those answers too, not just from training.
Do all AI crawlers respect robots.txt?
The major named ones publicly commit to it, but compliance is voluntary and unverifiable from your side. Server-side user-agent blocking or WAF rules are the only enforcement — and user-agents can be spoofed.
Which directive wins when a bot matches two groups?
The most specific user-agent group, and only that one — a crawler matching a named group ignores the `*` group entirely. So adding `GPTBot: Disallow: /private` while `*` allows everything means GPTBot follows only its own rules, not both sets combined.
Does a paywall or login already keep content out?
From crawling, yes — a bot without credentials receives the same page an anonymous visitor does. What it does not cover is content already syndicated, quoted or mirrored elsewhere, which is why access control and licensing are separate questions.

Common errors and gotchas

  • Assuming a block is enforced, since robots.txt is voluntary and only well-behaved crawlers honour it.
  • Blocking a search crawler by accident with an over-broad user-agent pattern.
  • Using a stale user-agent name, since these are added and renamed frequently.
  • Expecting a block to remove content already used for training, which it cannot do.
  • Confusing crawling for search with crawling for training, which some operators do with separate agents.

Related Web & SEO tools

Private & free — this tool runs entirely in your browser.

NamecheapRegister a domain for your next project — from $1.98/yr.affiliate