BusinessMCP

AI crawler user-agent identifier

Paste a user-agent from your logs. Find out which bot it is, who runs it, whether it trains a model or fetches a page for a waiting reader, and the exact robots.txt line that governs it.

Runs entirely in your browser — nothing is uploaded.

Every AI crawler, and what it wants

The bots this tool identifies, with the robots.txt token for each. Two of them are not crawlers at all — they are tokens only, which is the most common source of confusion in this list.

BotOperatorPurposerobots.txt
AmazonbotAmazonIndexes for an answer engineAmazonbot
AnthropicAnthropicTrains a modelanthropic-ai
Claude-SearchBotAnthropicIndexes for an answer engineClaude-SearchBot
Claude-UserAnthropicFetches for a readerClaude-User
Claude-WebAnthropicFetches for a readerClaude-Web
ClaudeBotAnthropicTrains a modelClaudeBot
Applebot-ExtendedApplerobots.txt token onlyApplebot-Extended
BytespiderByteDanceTrains a modelBytespider
CohereCohereTrains a modelcohere-ai
CCBotCommon CrawlOpen web archiveCCBot
Google-ExtendedGooglerobots.txt token onlyGoogle-Extended
Meta-AIMetaTrains a modelmeta-externalagent
Meta-ExternalFetcherMetaFetches for a readermeta-externalfetcher
Meta-WebIndexerMetaIndexes for an answer enginemeta-webindexer
ChatGPT-UserOpenAIFetches for a readerChatGPT-User
GPTBotOpenAITrains a modelGPTBot
OAI-AdsBotOpenAIIndexes for an answer engineOAI-AdsBot
OAI-SearchBotOpenAIIndexes for an answer engineOAI-SearchBot
Perplexity-UserPerplexityFetches for a readerPerplexity-User
PerplexityBotPerplexityIndexes for an answer enginePerplexityBot

Knowing the bot is step one. Seeing them is step two.

Most AI crawlers fetch your HTML without running any JavaScript, so a normal analytics script never sees them — on our own site the gap was roughly 3,800 crawler requests a day against about 140 that were visible to the beacon. BusinessMCP captures them server-side and reports them by the same taxonomy this tool uses, next to whether those crawls turn into citations in ChatGPT, Gemini and Perplexity.

Frequently asked questions

How do I tell an AI crawler from a real visitor?

By the user-agent, which is what this tool reads — but only as a first pass. A user-agent is self-reported, so anything can claim to be Chrome, and plenty of automation does. The strings that matter are the declared ones: GPTBot, ClaudeBot, PerplexityBot and their peers identify themselves honestly because they want to be allowed, and those are the ones a robots.txt rule can actually govern.

What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

All three are OpenAI and they do opposite jobs. GPTBot crawls to train models. OAI-SearchBot indexes so ChatGPT can surface and link you. ChatGPT-User fetches one page because a person asked for it — a real reader is waiting. Blocking all three with one rule is common and usually not what the site owner meant: it opts out of training and also removes you from ChatGPT search.

Why does Google-Extended never show up in my logs?

Because it is not a crawler and is never sent as a user-agent. Googlebot does the fetching; Google-Extended is a robots.txt token that controls whether what Googlebot fetched may train Gemini and Vertex AI. Applebot-Extended works the same way for Apple Intelligence. Blocking either one does not affect your Search or Siri visibility.

Should I block AI crawlers?

That is a business decision and we are not going to make it for you, which is why nothing on this page colours a bot as good or bad. The useful frame: training crawlers take content and give nothing back directly, while answer-engine indexers are how you get cited and linked in ChatGPT and Perplexity. Most teams allow the indexers and decide separately about training.

How do I actually block one?

Add the exact token to robots.txt — the tool prints the right line for each bot. Note that robots.txt is a request, not an enforcement mechanism: the major declared crawlers honour it, and anything determined to ignore it needs a WAF rule instead. Our AI crawler checker tests a live domain to confirm what your current rules really allow.

How do I see which AI bots are hitting my site?

Server logs, or an analytics tool that keeps bot traffic instead of discarding it. Most JavaScript analytics never sees these bots at all, because they fetch HTML without executing scripts. BusinessMCP captures them server-side and reports them by the same taxonomy this tool uses.