Is your robots.txt blocking AI crawlers? GPTBot, ClaudeBot and PerplexityBot explained
Many sites block AI crawlers by accident, often through a blanket rule or a security setting. If you want to be recommended, you need to know which bots to allow and why.
By the Citeably team · · 5 min read
Why this matters
Assistants that search the web fetch pages with their own crawlers. If your robots.txt, firewall, or CDN blocks those crawlers, your pages can't be read or quoted, and you become much less likely to be recommended.
Companies now run separate crawlers for separate jobs, so you can often welcome the ones that power search and answers while opting out of ones used to train models. Check each company's current documentation, because these names and behaviours change.
The main AI crawlers
As documented by each company at the time of writing:
- GPTBot (OpenAI): collects content that may be used to train models.
- OAI-SearchBot (OpenAI): finds pages to show in ChatGPT's search results.
- ChatGPT-User (OpenAI): fetches a page when a person asks ChatGPT to open it.
- ClaudeBot (Anthropic): general crawler that may collect training data.
- Claude-SearchBot and Claude-User (Anthropic): support search quality and user-requested page visits.
- PerplexityBot and Perplexity-User (Perplexity): index pages for Perplexity's answers and fetch pages on a person's request.
- Google-Extended (Google): a robots.txt control for some uses of your content by Google's Gemini products. It is separate from Google Search.
- Applebot-Extended (Apple) and CCBot (Common Crawl): used for AI training and open web datasets.
A robots.txt that welcomes AI search
If your goal is to be found and recommended, allow the search and user-triggered bots at minimum. This example allows the main ones and keeps private areas blocked:
User-agent: * Allow: / Disallow: /admin Disallow: /api/ User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User Allow: / Disallow: /admin Disallow: /api/ Sitemap: https://yourdomain.com/sitemap.xml
The trade-off
Allowing training crawlers (such as GPTBot or ClaudeBot) can make it more likely that models know about your brand without searching. Blocking them keeps your content out of future training. That's a business decision. Publishers often block training bots; brands that want recall usually allow them. Whatever you choose, make it deliberate rather than accidental.
How to check for accidental blocks
- Open yourdomain.com/robots.txt and look for “Disallow: /” under a wildcard or a named bot.
- Check your CDN or firewall (for example bot-protection modes) for rules that block unknown automated traffic.
- Make sure key pages work without JavaScript, because many crawlers don't run scripts.
- Run Citeably's free check: it tests your robots.txt against the major AI crawlers and flags anything blocked.
Frequently asked questions
Will allowing AI crawlers get me recommended?
Not by itself, but blocking them can keep you out entirely. Being crawlable is a necessary first step, not a guarantee.
Does blocking GPTBot hide me from ChatGPT search?
Not necessarily: OpenAI documents separate crawlers for training and for search. Check OpenAI's current documentation and test your setup.