Every bot that decides your AI visibility
Whether ChatGPT, Claude, Perplexity, or Google can read — and recommend — your business comes down to a handful of crawlers. What each one does, and whether you should let it in.
AI search crawlers
OAI-SearchBot
Builds the index behind ChatGPT search — this is the bot that decides if ChatGPT can cite you.
PerplexityPerplexityBot
Indexes pages for Perplexity's answer engine — Perplexity cites sources on almost every answer.
AmazonAmazonbot
Feeds Alexa and Amazon's AI answers.
DuckDuckGoDuckAssistBot
Powers DuckDuckGo's AI-assisted answers.
On-demand AI fetchers
AI training crawlers
GPTBot
Collects pages to train ChatGPT's underlying models.
AnthropicClaudeBot
Crawls pages for Claude's models and search index.
GoogleGoogle-Extended
Controls whether Google may train Gemini on your content. Blocking it does NOT affect Google Search or AI Overviews.
Common CrawlCCBot
Builds a public web archive that many AI companies train models on. Blocking it quietly removes you from future models' knowledge.
ByteDanceBytespider
Trains ByteDance's models (Doubao).
AppleApplebot-Extended
Controls whether Apple may use your content for Apple Intelligence.
MetaMeta-ExternalAgent
Collects pages to train Meta's Llama models.
Search crawlers
Which of these can actually reach your site? The free SeeGeo audit evaluates your robots.txt against every crawler on this page — plus your CDN and JavaScript rendering.
Run a free auditFrequently asked questions
How can I check whether AI crawlers like GPTBot can read my website?
Check two things: that your robots.txt does not disallow the crawler, and that your server or CDN actually returns the page when a request arrives with that crawler's user agent. The second check matters because firewalls such as Cloudflare can block AI bots regardless of robots.txt. SeeGeo's free crawler check does both for GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot in a few seconds.
Which AI crawlers should a small business allow?
Allow the search-type crawlers that feed answers customers see — OAI-SearchBot and ChatGPT-User for ChatGPT, Googlebot for Google's AI Overviews, PerplexityBot, and ClaudeBot — since blocking them removes you from recommendations. The training-only crawlers (GPTBot, Google-Extended, CCBot) are a genuine choice: allowing them helps future models know you exist, blocking them keeps your content out of training data.
Does robots.txt alone control AI crawlers?
No. robots.txt is a request that well-behaved crawlers honour, but your CDN or firewall can block a crawler that robots.txt allows, and a crawler can be listed as allowed while your pages render only with JavaScript it never runs. Access is only real when the fetch succeeds and the content is in the HTML.