Crawler directory

Every bot that reads your site, and what it wants.

66 AI crawlers, AI assistants and search bots from 16 companies. Who runs each one, whether it follows robots.txt, and how to block, allow or verify it, taken from each company's own documentation.

OpenAI Anthropic Google Microsoft Perplexity Apple Meta Amazon Mistral AI Common Crawl Yandex Ahrefs Semrush Baidu Majestic Slack

AI assistants

Open a page because a person asked a chatbot about it.

AI search

Index pages so AI answers can cite and link them.

AI training

Collect pages that may be used to train AI models.

Search engines

Crawl pages for classic search results.

Other bots

Check ads, build link previews and power SEO tools.

OpenAI OAI-AdsBot Checks that a landing page submitted as a ChatGPT ad is safe. Ignores robots.txt Google Google-CloudVertexBot Crawls a site when its owner asks Vertex AI to build an agent from it. Follows robots.txt Google Google-InspectionTool Fetches a page when someone runs the Rich Results Test or URL Inspection in Search Console. Follows robots.txt Google AdsBot-Google Checks the quality of landing pages used in Google Ads. Follows robots.txt Google AdsBot-Google-Mobile Checks the quality of mobile landing pages used in Google Ads. Follows robots.txt Google Mediapartners-Google Reads pages that show AdSense or Ad Manager ads so the ads match the content. Follows robots.txt Google APIs-Google Delivers push notification messages sent through Google APIs. Follows robots.txt Google Google-Safety Crawls for abuse checks, such as finding malware. Ignores robots.txt Google FeedFetcher-Google Fetches RSS and Atom feeds for Google News and WebSub. May skip robots.txt Google GoogleProducer Processes the feeds publishers supply to Google News through Publisher Center. May skip robots.txt Google Google-Read-Aloud Fetches a page so Google can read it out loud with text to speech. May skip robots.txt Google GoogleMessages Builds the link preview when someone sends your page in Google Messages. May skip robots.txt Google Google-Pinpoint Fetches a page that a Pinpoint user added as a source to their research collection. May skip robots.txt Google Google-Site-Verification Fetches the token that proves you own a site in Search Console. May skip robots.txt Google Google-CWS Requests the URLs a developer lists in a Chrome Web Store extension or theme. May skip robots.txt Meta facebookexternalhit Builds the link preview when someone shares your page in one of Meta's apps. Follows robots.txt Meta meta-externalads Crawls pages to improve Meta's advertising and business products. Follows robots.txt Ahrefs AhrefsBot Crawls the web for Ahrefs' SEO and backlink data and for the Yep search engine. Follows robots.txt Ahrefs AhrefsSiteAudit Runs the Ahrefs Site Audit that a site owner starts on their own site. Follows robots.txt Semrush SemrushBot Crawls links for Semrush's backlink index. Follows robots.txt Semrush SiteAuditBot Crawls a site for Semrush's technical SEO audit. Follows robots.txt Semrush SemrushBot-BA Checks links for Semrush Backlink Audit. Follows robots.txt Semrush SemrushBot-SI Fetches pages for the Semrush On Page SEO Checker. Follows robots.txt Semrush SemrushBot-SWA Checks whether URLs are reachable for the Semrush SEO Writing Assistant. Follows robots.txt Semrush SplitSignalBot Fetches pages for SplitSignal, Semrush's SEO A/B testing tool. Follows robots.txt Majestic MJ12bot Crawls the web to map links between sites for Majestic's link index. Follows robots.txt Slack Slackbot-LinkExpanding Reads a page's meta tags to show a preview when someone posts its link in Slack. Ignores robots.txt Slack Slack-ImgProxy Caches images posted in Slack and serves them over HTTPS. Ignores robots.txt Slack Slackbot Handles Slack API requests and outgoing webhooks that the other Slack bots don't cover. Ignores robots.txt

Start from this robots.txt.

It blocks the crawlers that train models and keeps the ones that send you readers. Adjust the lists with the pages above.

  • Block AI training

    Training crawlers send no visitors back. Disallow them if you don't want your pages used to train models.

  • Keep AI search and assistants

    They decide whether AI answers cite and link you, and assistant fetchers visit because a person asked about your page.

  • Verify by IP, not by name

    Any script can copy a user agent. Check the IP against the company's published list, or look up its reverse DNS.

robots.txt

# Block AI training

User-agent: GPTBot

User-agent: ClaudeBot

User-agent: CCBot

User-agent: Google-Extended

User-agent: Applebot-Extended

User-agent: meta-externalagent

Disallow: /

# Keep AI search and assistants

User-agent: OAI-SearchBot

User-agent: ChatGPT-User

User-agent: Claude-SearchBot

User-agent: Claude-User

User-agent: PerplexityBot

Allow: /

$ host 18.97.14.84

18-97-14-84.crawl.commoncrawl.org

Questions, answered.

An AI crawler is a bot an AI company runs to read web pages. Some collect pages to train models, some build the index an AI assistant searches, and some fetch a page because a person asked about it. Each kind wants something different, so each needs its own rule.

Block the training crawlers if you don't want your content used to train models; they send no visitors back. Keep AI search crawlers and assistant fetchers allowed if you want to be cited in AI answers, since those visits come from real people asking about your topic.

No. OpenAI uses GPTBot to collect training data and OAI-SearchBot for ChatGPT search, and each reads its own robots.txt rules. You can block GPTBot and still allow OAI-SearchBot.

It works for the crawlers that follow it, which is most of the ones listed here. Some assistant fetchers may skip it because a person started the request, and a bot that fakes its user agent ignores it. For those, block the request at your server or firewall.

An AI training crawler, like GPTBot or ClaudeBot, collects pages that may train future models and sends no visitors. An AI search crawler, like OAI-SearchBot or PerplexityBot, builds the index an assistant searches before it answers, which is how your pages get cited and linked.

If you want to be found, allow the AI search crawlers and the assistant fetchers, and decide on training crawlers separately. Allowing OAI-SearchBot, Claude-SearchBot and PerplexityBot keeps you in the answers of ChatGPT, Claude and Perplexity.

List each AI crawler in robots.txt under its own User-agent line with "Disallow: /", or build the file with our free robots.txt generator. Then block the fetchers that may skip robots.txt at your firewall.

Check the IP, not the user agent. Most companies publish their crawlers' IP ranges or a reverse DNS name, and each crawler page here says which one applies. The free bot verifier runs the check for you.

Most AI crawlers never run JavaScript, so an analytics script in the page never sees them. They show up in server logs, or with a check that runs on the server before the page loads, such as the NoirTrack server SDK.

Still have questions? We're happy to help.

See which bots really read your site.

NoirTrack counts AI crawlers and search bots apart from real visitors, and its firewall can block them before they reach your pages.

Start free trial

Free 14-day trial, no credit card. The server SDK also sees bots that never run JavaScript.