How to verify a bot
- Find the request in your logs. Copy its user agent string and the client IP address.
- Paste both into the verifier above and run it.
- Read the verdict. "Verified" means the IP belongs to the company the user agent names. "Fake" means it doesn't. "Can't be verified" means the company publishes no way to check.
The two example buttons show a real Googlebot and a fake GPTBot, so you can see both answers before using your own logs.
Why a user agent proves nothing
A user agent is a line of text the client chooses. Any script can send Mozilla/5.0 (compatible; Googlebot/2.1) with one header. Scrapers do exactly that, because many sites wave through anything that looks like Google or a well-known AI crawler.
The IP address is different: it is where the request came from, and it can't be faked without the traffic failing to reach the scraper. So every verification method checks the IP.
The two ways to verify a crawler
1. The company's published IP list. OpenAI, Anthropic, Google, Microsoft, Apple, Perplexity, Amazon, Mistral and Common Crawl publish the IP ranges their crawlers use. If the IP falls inside one of those ranges, the request is real.
2. Reverse DNS, confirmed forward. Look up the hostname for the IP, check that it belongs to the company, then look that hostname up again and make sure it returns the same IP:
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1
The second step matters: anyone who controls an IP can set its reverse DNS to say googlebot.com, but only Google can make googlebot.com point back to it.
Who publishes what
| Company | Crawlers | How to verify |
|---|---|---|
| Googlebot, Google-Extended and others | IP list and reverse DNS (googlebot.com, google.com) |
|
| OpenAI | GPTBot, OAI-SearchBot, ChatGPT-User | IP lists |
| Anthropic | ClaudeBot, Claude-SearchBot, Claude-User | IP list |
| Microsoft | bingbot | IP list |
| Apple | Applebot | IP list and reverse DNS (applebot.apple.com) |
| Perplexity | PerplexityBot, Perplexity-User | IP lists |
| Common Crawl | CCBot | IP list and reverse DNS (crawl.commoncrawl.org) |
| Yandex | YandexBot and others | Reverse DNS (yandex.ru, yandex.net, yandex.com) |
| Meta, Semrush | meta-externalagent, SemrushBot and others | Neither, so only the user agent |
Every crawler's page in the crawler directory links the company's own documentation and list.
How to find the user agent and IP in your logs
A typical Nginx or Apache access log line holds both:
66.249.66.1 - - [06/Oct/2026:10:12:01 +0000] "GET /pricing HTTP/1.1" 200 5123 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
The first field is the IP and the last quoted field is the user agent. Behind a CDN or proxy, the first field is the proxy's address; the visitor's real IP is in a header such as X-Forwarded-For or your CDN's own header. Only trust those headers when they come from your proxy.
What to do about fake bots
- Block them at the server or firewall. robots.txt won't help: a bot that lies about its name ignores the file too.
- Watch for patterns. Fake crawlers usually come from cloud and hosting networks and request pages far faster than the real crawler would.
- Let the real ones through. Blocking every request that claims to be Googlebot hurts your search rankings. Verify, then decide.
The NoirTrack firewall blocks bots, scrapers and datacenter traffic before they reach your pages, and the server SDK sees crawlers that never run JavaScript, so you can see which bots visit and how often.