How to use the robots.txt generator
- Pick a starting point. "Block AI training" keeps AI search and assistants allowed, "Block all AI" closes every AI crawler, and "Allow everything" opens the site to all.
- Adjust the groups. Switch AI training, AI search, AI assistants or search engines to Allow or Block. To set crawlers one by one, open the full list and tick the ones to block.
- Add your sitemap if you have one, so every crawler that reads the file can find it.
- Copy or download the file and upload it to the root of your domain, so it loads at
https://yourdomain.com/robots.txt.
Everything runs in your browser. Nothing you pick is sent anywhere.
What a robots.txt file looks like
# Keep AI training crawlers out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
Disallow: /
# Everyone else may read the whole site
User-agent: *
Allow: /
Sitemap: https://yourdomain.com/sitemap.xml
Lines starting with # are comments. Each group starts with one or more User-agent lines and lists the rules for those crawlers.
robots.txt rules explained
| Line | What it does |
|---|---|
User-agent: GPTBot |
Starts a group for one crawler. Several lines in a row share the same rules. |
User-agent: * |
The group for every crawler that has no group of its own. |
Disallow: /path |
Keeps the crawler out of that path and everything under it. Disallow: / closes the whole site. |
Allow: /path |
Opens a path inside a closed folder. When Allow and Disallow match equally, Allow wins. |
Sitemap: URL |
Points crawlers to your sitemap. It applies to every crawler, wherever it sits in the file. |
Crawl-delay: 10 |
Asks for seconds between requests. Bing reads it; Google ignores it. |
Paths can use * for "anything" and end with $ for "ends here", so Disallow: /*.pdf$ blocks every PDF.
Ready-made setups
Stay in AI answers, stay out of training. The most common choice for businesses that want to be found:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /
User-agent: *
Allow: /
Block every AI crawler, keep search engines. Add the AI search crawlers and assistant fetchers to the list above, such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User.
Close one folder for everyone.
User-agent: *
Disallow: /admin/
Google-Extended and Applebot-Extended never crawl anything. They are switches that Google and Apple read to decide whether pages their other crawlers fetch may train their models. Blocking them keeps you in Google and Apple search.
How to test your robots.txt
- Open
https://yourdomain.com/robots.txtin a browser. You should see the plain text file, not a web page. - Run the AI crawler checker on your domain to see what every crawler makes of it.
- In Google Search Console, the robots.txt report shows the version Google last fetched and any lines it couldn't read.
What robots.txt can't do
- It isn't security. The file is public, so anyone can read the paths you list. Protect private pages with a login, not a
Disallow. - It doesn't remove pages from search. A blocked page can still be indexed if other sites link to it. To keep a page out of results, let it be crawled and add a
noindextag. - It only works on bots that follow it. Scrapers and fake crawlers ignore it. Check a suspicious request with the bot verifier, and block bad traffic with a firewall such as the NoirTrack firewall.