A new kind of visitor has become one of the busiest on the web: AI crawlers. They read your pages to train language models, to build search indexes for AI assistants, or to fetch a page on the spot because someone asked a chatbot a question.
Some of that traffic can send you customers. Some of it just copies your content. The trick is telling them apart.
The three kinds of AI bots
Not all AI bots want the same thing. It helps to sort them into three groups:
- Training crawlers collect pages to train future models. Examples: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Google-Extended. They read a lot and send no visitors back.
- AI search crawlers build the index an AI assistant searches when answering. Examples: OAI-SearchBot, PerplexityBot. Being in that index is how you get cited in answers.
- AI agents fetch a page right now because a person asked. Examples: ChatGPT-User, Claude-User. Behind each visit is a real person who wanted your page.
What blocking costs and what it saves
| Block | Allow | |
|---|---|---|
| Training crawlers | Your content is not used for training. Server load drops. | Your brand may be better known to models. No direct visitors. |
| AI search crawlers | You will not be cited in AI answers. | You can be quoted and linked in answers, which sends real visitors. |
| AI agents | The person who asked gets nothing from you. | A real person reads your page through the assistant. |
For most businesses that want to be found, a sensible default is: block training crawlers if you care about your content being copied, allow AI search crawlers and AI agents. Being cited in AI answers is quickly becoming as important as ranking on Google.
How to block or allow them
There are two ways, and they work best together.
robots.txt is a polite request. Well-behaved bots like GPTBot respect it:
User-agent: GPTBot
Disallow: /
User-agent: CCBot
Disallow: /
A firewall enforces it. Bots that ignore robots.txt, or pretend to be a normal browser, are turned away anyway. In NoirTrack the firewall has an AI scrapers switch that blocks the training crawlers, and a separate AI agents switch that lets assistants through when a person asks. Search engines stay allowed by default.
Measure before you decide
Before blocking anything, look at what is actually visiting. The Bots & security panel in NoirTrack splits bot traffic into categories such as AI answers, training and indexing, and shows which pages each bot reads. One thing to know: bots that do not run JavaScript, which includes most AI crawlers, are only seen by a server-side check. Add the server SDK if you want the full picture.
If you find that AI answer engines visit your pricing and docs pages often, that is a good sign: you are probably being cited. Keep those doors open.