How to check which AI crawlers can read your site
- Enter your domain in the box above, like
yoursite.com. A full URL works too; the checker keeps only the domain. - Run the check. The tool fetches
/robots.txtand/llms.txtfrom that domain, the same files a crawler reads before anything else. - Read the verdicts. Every AI assistant, AI search, AI training and search engine crawler in our crawler directory gets one: allowed or blocked, plus whether your file names it directly.
- Fix what you don't like. Open the robots.txt generator, pick the crawlers you want out, and replace your file.
The check takes a few seconds. The result is kept for ten minutes and has its own link, so you can send it to a teammate or a client.
How to read the results
| What you see | What it means |
|---|---|
| Blocked | Your robots.txt keeps this crawler out of the whole site. |
| Allowed | The crawler may read your public pages. |
| Some paths blocked | The crawler is allowed, but some folders (often /admin or /cart) are closed to it. |
| Named in robots.txt | Your file has a User-agent group just for this crawler. Everything else falls back to User-agent: *. |
| robots.txt missing | No file was found, so every crawler may read every page. |
A warning appears when a block costs you something: AI search crawlers that can't read your pages can't cite them, and assistant fetchers only visit because a person asked about your page.
The four kinds of crawlers we check
| Kind | Examples | What blocking it does |
|---|---|---|
| AI training | GPTBot, ClaudeBot, CCBot, Google-Extended | Keeps your pages out of future training data. Sends no visitors either way. |
| AI search | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Removes your pages from the index those assistants cite in answers. |
| AI assistants | ChatGPT-User, Claude-User, Perplexity-User | Turns away a person who asked an assistant about your page. |
| Search engines | Googlebot, bingbot, Applebot, YandexBot | Removes your pages from that engine's results. |
Most sites that want to stay findable block the first row and allow the other three. The crawler directory explains every bot, with links to the company's own documentation.
How robots.txt decides who gets in
Crawlers read robots.txt the way RFC 9309 describes, and the checker does the same:
- Groups. Rules sit under one or more
User-agentlines. A crawler uses the group that names it, and only falls back toUser-agent: *when no group names it. - Longest match wins. If both an
Allowand aDisallowrule match a path, the longer rule decides. On a tie,Allowwins. - Empty means open.
Disallow:with nothing after it blocks nothing.
Here is a small file and what four crawlers make of it:
User-agent: GPTBot
Disallow: /
User-agent: *
Disallow: /admin
| Crawler | Group it uses | Result |
|---|---|---|
| GPTBot | its own | Blocked |
| OAI-SearchBot | * |
Allowed, /admin closed |
| ClaudeBot | * |
Allowed, /admin closed |
| Googlebot | * |
Allowed, /admin closed |
Notice that blocking GPTBot does nothing to OAI-SearchBot. OpenAI runs them as separate crawlers on purpose, so you can stay out of training and still appear in ChatGPT search.
Common mistakes that block AI search by accident
- A blanket
Disallow: /underUser-agent: *. Often left over from a staging site. It blocks every crawler that has no group of its own, AI search and search engines included. - Blocking the wrong OpenAI bot. People add
User-agent: ChatGPT-UserorOAI-SearchBotmeaning to stop training. The training crawler isGPTBot. - A CMS or host setting. "Discourage search engines" switches and one-click "block AI bots" options can rewrite robots.txt or block at the edge without you noticing.
- A firewall that disagrees with the file. robots.txt can allow a crawler that your CDN or server then blocks. This checker reads the file only, so a firewall block won't show here.
What to do after the check
- Want different rules? Build a new file with the robots.txt generator.
- Wondering whether a crawler in your logs is real? Paste it into the bot verifier.
- Want to see which crawlers actually visit, and how often? Most AI crawlers never run JavaScript, so page-based analytics miss them. The NoirTrack server SDK sees every request, and the firewall can block the ones you choose.