Crawler directory
Every bot that reads your site, and what it wants.
66 AI crawlers, AI assistants and search bots from 16 companies. Who runs each one, whether it follows robots.txt, and how to block, allow or verify it, taken from each company's own documentation.
AI assistants
Open a page because a person asked a chatbot about it.
OpenAI
ChatGPT-User
Opens a page when someone using ChatGPT or a custom GPT asks about it.
May skip robots.txt
Anthropic
Claude-User
Fetches a page when a person asks Claude something that needs it.
Follows robots.txt
Google
Google-Agent
Browses the web and takes actions on pages when a person asks a Google AI agent to.
May skip robots.txt
Google
Google-GeminiNotebook
Fetches a page that a person added as a source in a Gemini Notebook project.
May skip robots.txt
Perplexity
Perplexity-User
Visits a page when a Perplexity user asks a question that needs it.
May skip robots.txt
Meta
meta-externalfetcher
Fetches a link when a person asks Meta AI to open it.
May skip robots.txt
Amazon
Amzn-User
Fetches up-to-date pages for a person's request, such as an Alexa question.
May skip robots.txt
Mistral AI
MistralAI-User
Visits a page to answer a person's question in Mistral's assistant and cite it.
Follows robots.txt
AI search
Index pages so AI answers can cite and link them.
OpenAI
OAI-SearchBot
Indexes pages so they can appear and be linked in ChatGPT search answers.
Follows robots.txt
Anthropic
Claude-SearchBot
Crawls pages to improve the search results Claude cites.
Follows robots.txt
Perplexity
PerplexityBot
Indexes pages so Perplexity can surface and link them in its answers.
Follows robots.txt
Apple
Applebot
Crawls pages for search in Spotlight, Siri and Safari; Apple may also use what it collects to train its foundation models.
Follows robots.txt
Meta
meta-webindexer
Reads pages to improve the search results in Meta AI.
Follows robots.txt
Amazon
Amzn-SearchBot
Crawls pages for search in Amazon products. Amazon says it is not used for AI training.
robots.txt not stated
Mistral AI
MistralAI-Index
Crawls pages to build the index behind Mistral's search.
Follows robots.txt
AI training
Collect pages that may be used to train AI models.
OpenAI
GPTBot
Collects public pages that OpenAI may use to train its foundation models.
Follows robots.txt
Anthropic
ClaudeBot
Collects public web content that Anthropic may use to train and improve Claude.
Follows robots.txt
Google
Google-Extended
A robots.txt switch that decides whether Google may use your pages to train Gemini models and to ground their answers.
robots.txt switch, never crawls
Apple
Applebot-Extended
A robots.txt switch that opts your pages out of training Apple's foundation models.
robots.txt switch, never crawls
Meta
meta-externalagent
Crawls pages to train Meta's AI models and to improve its products.
Follows robots.txt
Amazon
Amazonbot
Crawls pages to improve Amazon products and to train Amazon's AI models.
Follows robots.txt
Mistral AI
MistralAI-Training
Collects web content for datasets that train Mistral's models.
Follows robots.txt
Common Crawl
CCBot
Builds Common Crawl's open archive of the web, which many AI labs use as training data.
Follows robots.txt
Search engines
Crawl pages for classic search results.
Google
Googlebot
Crawls pages for Google Search, Discover, Images, Video and News.
Follows robots.txt
Google
GoogleOther
A general crawler Google product teams use to fetch public content outside Search.
Follows robots.txt
Google
Storebot-Google
Crawls product pages for Google Shopping.
Follows robots.txt
Google
Googlebot-Image
Crawls images for Google Images, Discover and the image features in Search.
Follows robots.txt
Google
Googlebot-Video
Crawls videos for the video features in Google Search.
Follows robots.txt
Google
Googlebot-News
A robots.txt name for Google News: it crawls with the regular Googlebot user agents, and these rules decide what appears in News.
Follows robots.txt
Google
GoogleOther-Image
The image version of GoogleOther, fetching public image URLs for Google product teams.
Follows robots.txt
Google
GoogleOther-Video
The video version of GoogleOther, fetching public video URLs for Google product teams.
Follows robots.txt
Microsoft
bingbot
Crawls pages for Bing search.
Follows robots.txt
Yandex
YandexBot
Crawls pages for Yandex search.
Follows robots.txt
Yandex
YandexImages
Crawls images for Yandex Images.
Follows robots.txt
Yandex
YandexVideo
Crawls videos for Yandex video search.
Follows robots.txt
Yandex
YandexMobileBot
Checks whether a page is built for mobile devices.
Follows robots.txt
Baidu
Baiduspider
Crawls pages for Baidu search.
Follows robots.txt
Other bots
Check ads, build link previews and power SEO tools.
OpenAI
OAI-AdsBot
Checks that a landing page submitted as a ChatGPT ad is safe.
Ignores robots.txt
Google
Google-CloudVertexBot
Crawls a site when its owner asks Vertex AI to build an agent from it.
Follows robots.txt
Google
Google-InspectionTool
Fetches a page when someone runs the Rich Results Test or URL Inspection in Search Console.
Follows robots.txt
Google
AdsBot-Google
Checks the quality of landing pages used in Google Ads.
Follows robots.txt
Google
AdsBot-Google-Mobile
Checks the quality of mobile landing pages used in Google Ads.
Follows robots.txt
Google
Mediapartners-Google
Reads pages that show AdSense or Ad Manager ads so the ads match the content.
Follows robots.txt
Google
APIs-Google
Delivers push notification messages sent through Google APIs.
Follows robots.txt
Google
Google-Safety
Crawls for abuse checks, such as finding malware.
Ignores robots.txt
Google
FeedFetcher-Google
Fetches RSS and Atom feeds for Google News and WebSub.
May skip robots.txt
Google
GoogleProducer
Processes the feeds publishers supply to Google News through Publisher Center.
May skip robots.txt
Google
Google-Read-Aloud
Fetches a page so Google can read it out loud with text to speech.
May skip robots.txt
Google
GoogleMessages
Builds the link preview when someone sends your page in Google Messages.
May skip robots.txt
Google
Google-Pinpoint
Fetches a page that a Pinpoint user added as a source to their research collection.
May skip robots.txt
Google
Google-Site-Verification
Fetches the token that proves you own a site in Search Console.
May skip robots.txt
Google
Google-CWS
Requests the URLs a developer lists in a Chrome Web Store extension or theme.
May skip robots.txt
Meta
facebookexternalhit
Builds the link preview when someone shares your page in one of Meta's apps.
Follows robots.txt
Meta
meta-externalads
Crawls pages to improve Meta's advertising and business products.
Follows robots.txt
Ahrefs
AhrefsBot
Crawls the web for Ahrefs' SEO and backlink data and for the Yep search engine.
Follows robots.txt
Ahrefs
AhrefsSiteAudit
Runs the Ahrefs Site Audit that a site owner starts on their own site.
Follows robots.txt
Semrush
SemrushBot
Crawls links for Semrush's backlink index.
Follows robots.txt
Semrush
SiteAuditBot
Crawls a site for Semrush's technical SEO audit.
Follows robots.txt
Semrush
SemrushBot-BA
Checks links for Semrush Backlink Audit.
Follows robots.txt
Semrush
SemrushBot-SI
Fetches pages for the Semrush On Page SEO Checker.
Follows robots.txt
Semrush
SemrushBot-SWA
Checks whether URLs are reachable for the Semrush SEO Writing Assistant.
Follows robots.txt
Semrush
SplitSignalBot
Fetches pages for SplitSignal, Semrush's SEO A/B testing tool.
Follows robots.txt
Majestic
MJ12bot
Crawls the web to map links between sites for Majestic's link index.
Follows robots.txt
Slack
Slackbot-LinkExpanding
Reads a page's meta tags to show a preview when someone posts its link in Slack.
Ignores robots.txt
Slack
Slack-ImgProxy
Caches images posted in Slack and serves them over HTTPS.
Ignores robots.txt
Slack
Slackbot
Handles Slack API requests and outgoing webhooks that the other Slack bots don't cover.
Ignores robots.txt
Start from this robots.txt.
It blocks the crawlers that train models and keeps the ones that send you readers. Adjust the lists with the pages above.
Block AI training
Training crawlers send no visitors back. Disallow them if you don't want your pages used to train models.
Keep AI search and assistants
They decide whether AI answers cite and link you, and assistant fetchers visit because a person asked about your page.
Verify by IP, not by name
Any script can copy a user agent. Check the IP against the company's published list, or look up its reverse DNS.
# Block AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /
# Keep AI search and assistants
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Allow: /
$ host 18.97.14.84
18-97-14-84.crawl.commoncrawl.org
Questions, answered.
An AI crawler is a bot an AI company runs to read web pages. Some collect pages to train models, some build the index an AI assistant searches, and some fetch a page because a person asked about it. Each kind wants something different, so each needs its own rule.
Block the training crawlers if you don't want your content used to train models; they send no visitors back. Keep AI search crawlers and assistant fetchers allowed if you want to be cited in AI answers, since those visits come from real people asking about your topic.
No. OpenAI uses GPTBot to collect training data and OAI-SearchBot for ChatGPT search, and each reads its own robots.txt rules. You can block GPTBot and still allow OAI-SearchBot.
It works for the crawlers that follow it, which is most of the ones listed here. Some assistant fetchers may skip it because a person started the request, and a bot that fakes its user agent ignores it. For those, block the request at your server or firewall.
An AI training crawler, like GPTBot or ClaudeBot, collects pages that may train future models and sends no visitors. An AI search crawler, like OAI-SearchBot or PerplexityBot, builds the index an assistant searches before it answers, which is how your pages get cited and linked.
If you want to be found, allow the AI search crawlers and the assistant fetchers, and decide on training crawlers separately. Allowing OAI-SearchBot, Claude-SearchBot and PerplexityBot keeps you in the answers of ChatGPT, Claude and Perplexity.
List each AI crawler in robots.txt under its own User-agent line with "Disallow: /", or build the file with our free robots.txt generator. Then block the fetchers that may skip robots.txt at your firewall.
Check the IP, not the user agent. Most companies publish their crawlers' IP ranges or a reverse DNS name, and each crawler page here says which one applies. The free bot verifier runs the check for you.
Most AI crawlers never run JavaScript, so an analytics script in the page never sees them. They show up in server logs, or with a check that runs on the server before the page loads, such as the NoirTrack server SDK.
Still have questions? We're happy to help.
See which bots really read your site.
NoirTrack counts AI crawlers and search bots apart from real visitors, and its firewall can block them before they reach your pages.
Start free trialFree 14-day trial, no credit card. The server SDK also sees bots that never run JavaScript.