Allow the crawlers that read your site to answer a live question — OAI-SearchBot, ChatGPT-User, PerplexityBot and Anthropic's search and on-demand fetchers — because those are the only bots that decide whether you appear in an AI answer. The training crawlers, GPTBot and ClaudeBot, are a separate and far lower-stakes call: OpenAI states that blocking GPTBot has no effect on whether ChatGPT search cites you, while blocking OAI-SearchBot removes you from those answers outright. And two names that sit on nearly every "block these AI bots" list — Google-Extended and Applebot-Extended — aren't crawlers at all.
We build an AI-visibility plugin for WordPress, and the mistake we see most is treating "AI bots" as one switch. It isn't one decision, it's three — and for a storefront that three-way call between block, charge and allow is worked into a decision matrix in the 2026 AI-crawler decision for a small WooCommerce store. Once you sort each bot by the job it does, the answer to "which do I allow" stops being a preference and becomes almost mechanical.
Which AI crawlers decide whether you appear in AI answers?
The search and user-triggered bots decide it — OAI-SearchBot and PerplexityBot build the indexes that AI answers are drawn from, and ChatGPT-User and Perplexity-User fetch a page the instant a person asks about it. Block any of these and you remove yourself from the exact moment an assistant would have cited you. These are the ones to allow without hesitation on any public site.
| Crawler | Run by | Job | Block it and you lose |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Builds ChatGPT's search/citation index | Your place in ChatGPT search answers |
| ChatGPT-User | OpenAI | Fetches a page when a user asks live | A visit a person requested in real time |
| PerplexityBot | Perplexity | Builds Perplexity's answer index | Your citations in Perplexity |
| Perplexity-User | Perplexity | User-triggered fetch during a query | The live, high-intent visit |
| ClaudeBot + Claude-User | Anthropic | Crawl plus on-demand fetch for Claude | Reachability for Claude's answers |
| GPTBot | OpenAI | Collects training data for future models | Only a slot in a future training set |
| Google-Extended | A robots.txt training-opt-out token — no crawler | Nothing about crawling or Search |
The user-triggered bots deserve special weight, because a request from ChatGPT-User or Perplexity-User means a human is sitting in a chat right now and asked the assistant to go read your page. That is the highest-intent visitor you will ever get, and it's the one most often swept up when someone blocks "AI" in a panic. Reachability for these bots is the first rung of the whole visibility ladder we lay out in why AI assistants never cite your blog — nothing you write matters until the bot that would quote it gets in.
Do the training crawlers actually affect your AI visibility?
No — blocking GPTBot or ClaudeBot's training crawl has no effect on whether you're cited today, because citations come from the search and user bots, not the training corpus. The only thing you forfeit by blocking a training crawler is a place in some future model's training data, and that is both unmeasurable and un-actionable: you can't verify you're in it, you can't tell what it did for you, and you can't tie a single visitor to it.
That asymmetry is the whole decision in one picture.
- 1
Block ChatGPT-User or Perplexity-User
Lose a visit a human asked for in real time — highest intent
- 2
Block OAI-SearchBot or PerplexityBot
Vanish from ChatGPT and Perplexity answers — full citation cost
- 3
Block ClaudeBot's fetchers
Lose reachability for Claude's answers
- 4
Block GPTBot / ClaudeBot training
Lose only a slot in a future training set — no live traffic
- 5
Block Google-Extended
Lose nothing about crawling — it isn't a crawler
Our honest position, because the field genuinely splits here: if you sell access to your content or hold rights you want to protect, blocking the training crawlers is a reasonable stance and costs you no citations. If you're a marketing site or a store that wants to be found, blocking them protects nothing of value and buys nothing you can measure. Judge it against traffic you can actually see — the referrals you can watch arriving from ChatGPT and Perplexity — not against a training benefit no one can confirm.
Why is Google-Extended the exception that isn't even a crawler?
Because Google-Extended is a control token, not a user-agent — it has no separate HTTP request, does no crawling of its own, and only tells Google whether it may use content it already fetched to train Gemini and its generative models. The crawling is still done by Googlebot under its own rules. Google introduced the split in 2023 precisely so a publisher could opt out of AI training while keeping full Search presence, and setting Google-Extended to Disallow leaves your rankings and your AI Overviews eligibility untouched.
Applebot-Extended works the same way for Apple — a training-use flag layered on top of the real crawler, Applebot, not a crawler you can block. So a site owner who proudly "blocked Google's and Apple's AI" has usually just recorded a training preference and changed nothing about who crawls them or where they rank. This is exactly the kind of setting that looks meaningful in an SEO plugin's toggle list and does something entirely different from what the label implies — the gap between naming and effect we get into in do you need an AI visibility tool if you already have Yoast or Rank Math.
Does allowing a bot in robots.txt guarantee it gets through?
No — a robots.txt Allow is an intention, and two things routinely overrule it. The first is enforcement at the network edge: a Cloudflare rule or firewall can return a 403 to a bot your robots.txt plainly welcomes, and because that block sits in front of WordPress, nothing inside your site records it — the failure mode we walk through in is Cloudflare silently blocking AI crawlers from your site.
The second is that the string only works on crawlers that honour it. In August 2025 Cloudflare documented Perplexity switching to an undeclared crawler that impersonated a normal Chrome browser — a generic Mozilla/5.0 ... Chrome/124.0.0.0 Safari/537.36 user-agent — rotating its networks and ignoring or skipping robots.txt once its declared PerplexityBot was blocked, and Cloudflare removed Perplexity from its verified-bot list over it. In the same tests ChatGPT-User fetched robots.txt, saw the block and stopped cleanly. The practical lesson: allowing the named bots is necessary but not sufficient — you confirm the outcome by fetching a live page as each bot and reading the status code, or by checking your server logs for the verified crawler IPs, never by trusting the robots.txt file alone. Of those two, the log check is the one that settles it, because a fetch carrying a borrowed user-agent can be blocked when the real crawler isn't — and cleared when it is.
So which should you allow — the short rule?
Allow every search and user bot, decide the training bots on your values, and stop worrying about the -Extended tokens because they don't crawl. For a public marketing site or store that wants AI traffic, the clean default is: allow OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User and Anthropic's Claude fetchers, and leave GPTBot and ClaudeBot allowed too unless you have a specific reason to protect your content from training.
Then verify, because the allow-list is where the job starts, not ends. Contexta's AI Visibility test fetches each page as the individually named bots and reports, per crawler, whether it got a 200 or hit a robots.txt rule or a Cloudflare edge block — so you see which of these bots actually reach your content rather than which ones your robots.txt says are welcome. A welcome sign on the door and an open door are not the same thing, and only one of them shows up in AI answers.
FAQ
Should I block GPTBot?
Only if you specifically want to keep your content out of future model training — blocking GPTBot has no effect on whether ChatGPT cites you. GPTBot is OpenAI's training crawler; the bot that decides your presence in ChatGPT's answers is OAI-SearchBot, which is a separate user-agent with separate rules. For a public marketing site or store, blocking GPTBot protects nothing of value and buys no measurable benefit, so most such sites should leave it allowed.
Does blocking Google-Extended hurt my Google rankings?
No — Google-Extended is a training-opt-out token, not a crawler, so blocking it leaves your Search rankings and AI Overviews eligibility untouched. It has no separate HTTP request of its own; the crawling is still done by Googlebot under its own robots.txt rules. Google introduced the split in 2023 precisely so publishers could opt out of Gemini training while keeping full Search presence, and Applebot-Extended works the same way for Apple.
What's the difference between OAI-SearchBot and ChatGPT-User?
OAI-SearchBot builds the index ChatGPT search draws from, while ChatGPT-User fetches a single page live when a person asks about it in a chat. The first is ongoing indexing; the second is on-demand retrieval triggered by a real user, which makes it your highest-intent AI visitor. Both are separate from GPTBot, OpenAI's training crawler, and allowing or blocking one has no effect on the others.
If I allow PerplexityBot, will Perplexity respect it?
Allowing PerplexityBot lets it crawl normally, but blocking is the case that isn't reliably respected — in August 2025 Cloudflare documented Perplexity switching to an undeclared crawler impersonating a normal Chrome browser after its declared bot was blocked, and de-listed it as a verified bot over it. So a robots.txt rule is an intention, not a guarantee. Confirm the real outcome by fetching a page as the bot and reading the status code, or by checking server logs for verified crawler IPs.
