Free Robots.txt Generator A robots.txt for the AI era.

Choose a starting point, then switch individual crawlers on or off. The file updates as you go. Copy it or download it and put it at the root of your domain.

How to use the robots.txt generator

  1. Pick a starting point. "Block AI training" keeps AI search and assistants allowed, "Block all AI" closes every AI crawler, and "Allow everything" opens the site to all.
  2. Adjust the groups. Switch AI training, AI search, AI assistants or search engines to Allow or Block. To set crawlers one by one, open the full list and tick the ones to block.
  3. Add your sitemap if you have one, so every crawler that reads the file can find it.
  4. Copy or download the file and upload it to the root of your domain, so it loads at https://yourdomain.com/robots.txt.

Everything runs in your browser. Nothing you pick is sent anywhere.

What a robots.txt file looks like

# Keep AI training crawlers out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
Disallow: /

# Everyone else may read the whole site
User-agent: *
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

Lines starting with # are comments. Each group starts with one or more User-agent lines and lists the rules for those crawlers.

robots.txt rules explained

Line What it does
User-agent: GPTBot Starts a group for one crawler. Several lines in a row share the same rules.
User-agent: * The group for every crawler that has no group of its own.
Disallow: /path Keeps the crawler out of that path and everything under it. Disallow: / closes the whole site.
Allow: /path Opens a path inside a closed folder. When Allow and Disallow match equally, Allow wins.
Sitemap: URL Points crawlers to your sitemap. It applies to every crawler, wherever it sits in the file.
Crawl-delay: 10 Asks for seconds between requests. Bing reads it; Google ignores it.

Paths can use * for "anything" and end with $ for "ends here", so Disallow: /*.pdf$ blocks every PDF.

Ready-made setups

Stay in AI answers, stay out of training. The most common choice for businesses that want to be found:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /

User-agent: *
Allow: /

Block every AI crawler, keep search engines. Add the AI search crawlers and assistant fetchers to the list above, such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User.

Close one folder for everyone.

User-agent: *
Disallow: /admin/

Google-Extended and Applebot-Extended never crawl anything. They are switches that Google and Apple read to decide whether pages their other crawlers fetch may train their models. Blocking them keeps you in Google and Apple search.

How to test your robots.txt

  • Open https://yourdomain.com/robots.txt in a browser. You should see the plain text file, not a web page.
  • Run the AI crawler checker on your domain to see what every crawler makes of it.
  • In Google Search Console, the robots.txt report shows the version Google last fetched and any lines it couldn't read.

What robots.txt can't do

  • It isn't security. The file is public, so anyone can read the paths you list. Protect private pages with a login, not a Disallow.
  • It doesn't remove pages from search. A blocked page can still be indexed if other sites link to it. To keep a page out of results, let it be crawled and add a noindex tag.
  • It only works on bots that follow it. Scrapers and fake crawlers ignore it. Check a suspicious request with the bot verifier, and block bad traffic with a firewall such as the NoirTrack firewall.

Questions, answered.

robots.txt is a plain text file at the root of a site that tells crawlers which pages they may fetch. Each group names a crawler with User-agent and lists Allow or Disallow rules for it.

At the root of your domain, so it loads at https://yourdomain.com/robots.txt. Crawlers only look there. Each subdomain needs its own file.

It is the group for every crawler that has no group of its own. A crawler named in its own group ignores the * rules, so blocking a bot by name works even if * allows everything.

Disallow keeps a crawler out of a path and Allow lets it in. When both match a page, the longer rule wins, and Allow wins a tie, so you can block a folder and still open one page inside it.

Block them if you don't want your content used to train models. They send no visitors, and blocking them doesn't affect search rankings or AI answers.

No. Google-Extended only controls whether Google may use your pages for Gemini training and grounding. Googlebot still crawls and ranks your site as before.

Yes, it helps. A Sitemap line with the full URL of your sitemap lets every crawler that reads robots.txt find it, including ones you never submitted it to.

No. Well-behaved crawlers follow it, some assistant fetchers may skip it because a person asked for the page, and bad bots ignore it. Use a firewall for those.

Try NoirTrack

Do this on every visit.

This tool checks one thing at a time. NoirTrack counts people, AI crawlers and bots on every visit, and blocks the ones you don't want.

Free 14-day trial, no credit card