Should You Block GPTBot? A Decision Guide
Whether to block GPTBot depends on what you sell. A publisher protecting paid content refuses every AI crawler. A business that wants to be recommended permits the retrieval crawlers — OAI-SearchBot, PerplexityBot, Claude-SearchBot — because blocking them removes citation eligibility entirely. Training crawlers are a separate decision. This page gives the agents, the directives and the verification test.
Before the first directive. Two things are required: write access to robots.txt at the domain root, and read access to the server access logs. robots.txt is a plain-text file at /robots.txt that names user-agents and states which paths each one is permitted to fetch. A user-agent is the identifying token an automated client sends with every request. A crawl directive is one Allow or Disallow line applied to the agents named above it. AI crawlers divide into two classes, with a third group beside them. A retrieval crawler fetches a page so an assistant can quote and link it in an answer. A training crawler copies text into a corpus used to build a later model version. A user-triggered fetcher retrieves one page because a person asked an assistant about it.
Questions
Common questions
Every answer below ships in the raw HTML, so an assistant reading this page without running a script still receives it.
Does blocking GPTBot remove a site from ChatGPT?
GPTBot is OpenAI's training crawler. ChatGPT search draws on OAI-SearchBot, a separate agent on a separate line. Refusing GPTBot while permitting OAI-SearchBot keeps a site citable in ChatGPT search answers and its text out of OpenAI's training corpus.
What happens to a site that refuses every AI crawler?
Every retrieval surface drops it. ChatGPT search, Perplexity and Claude search stop citing refused pages, and the text stays out of later model versions. Publishers accept the first outcome to secure the second.
Is anthropic-ai a real user agent?
Anthropic's crawler documentation names three agents: ClaudeBot, Claude-User and Claude-SearchBot. The token anthropic-ai circulates in published robots.txt files and third-party directories, and Anthropic does not list it. This page recommends no directive for it.
Related
Selecting a provider to set and re-check these directives is covered in how to choose an ai seo agency. The practice they belong to is defined in What Is Generative Engine Optimization?.
Notes and sources. Seven primary sources carry every token on this page, each checked 21 September 2026. OpenAI documents four agents — GPTBot for training generative AI foundation models, OAI-SearchBot for surfacing websites in ChatGPT search results, ChatGPT-User for pages visited when users ask ChatGPT questions, and OAI-AdsBot for ad landing-page safety validation — and states that sites opted out of OAI-SearchBot are not shown in ChatGPT search answers (OpenAI — Bots). Perplexity documents PerplexityBot as designed to surface and link websites in Perplexity search results, and Perplexity-User as a fetcher that robots.txt rules do not reliably govern because a user requested the page (Perplexity — Bots). Anthropic documents ClaudeBot as collecting web content that contributes to model training, Claude-SearchBot as navigating the web to improve search result quality, and Claude-User as accessing websites when individuals ask Claude questions, and states that its bots honour robots.txt directives (Anthropic — Does Anthropic crawl data from the web?). Google states that using Google-Extended does not affect a site's inclusion in Search and is not used as a ranking signal in Search (Google — Things to know about Google's web crawling), that Google-Extended has no separate HTTP request user agent string and operates in a control capacity (Google — Google common crawlers), that robots.txt directives for Googlebot are the control for how a site is crawled for Search including its AI features (Google — AI features and your website), and that only one group is valid for a particular crawler while other groups are ignored (Google — robots.txt specifications). Apple documents Applebot-Extended as an opt-out from training Apple's foundation models and states that pages disallowing it are still included in search results (Apple — About Applebot). Meta documents meta-externalagent as used for training foundation AI models or improving products by indexing content directly (Meta — Web crawlers). Common Crawl documents CCBot and the user-agent string CCBot/2.0 (https://commoncrawl.org/faq/) (Common Crawl — CCBot).
🔴 Two tokens carry no primary source. ByteDance publishes no reachable documentation for Bytespider; the reference URL inside its own user-agent string resolves inside China only, and no operator statement of robots.txt compliance was located on 21 September 2026. Anthropic's current documentation does not name anthropic-ai; the token appears only in third-party crawler directories and in copied robots.txt files. This page names both and recommends a directive for neither.
Next step
Find out which assistants name you today.
The audit runs 40 checks across five categories and queries seven assistants for your citation baseline. Findings in ten business days.