urltoolskit.org
URL utilities, in the browser
Say hi →

robots.txt generator

presets · AI crawler blocks · sitemap lines · linted as you type

A robots.txt is four directives and a great many ways to get them subtly wrong. Fill in the paths you want kept out of the crawl, tick the AI crawlers you would rather not feed, and take the file away — checked, as it is written, by the same rules the tester applies.

Block AI crawlers

Each of these is a separate group with Disallow: /. Hover a name to see who runs it.

Ready.

How to use

  1. Pick a preset if one matches, then edit it. The presets are starting points, not opinions — a shop that has no facets does not need the facet rules.
  2. Put one path per line under Disallow. A full URL pasted in is reduced to its path, and a missing leading slash is added, because both are the usual way this file gets broken.
  3. Use Allow for exceptions inside a disallowed path. Ordering does not matter: the longest matching rule wins and Allow breaks ties, so Allow: /admin/public/ beats Disallow: /admin/.
  4. Tick the AI crawlers you want to exclude. Each gets its own group, because a bot-specific group replaces the * group entirely rather than adding to it.
  5. Add your Sitemap URLs — absolute, with the scheme. This is the one line in the file that helps rather than restricts.
  6. Test a path against what you have built, then copy or download and put the file at https://yourdomain/robots.txt. Nowhere else: a subdirectory robots.txt is never read.

What a robots.txt does not do

It is not access control. The file is public, and listing /internal-reports/ in it tells everyone the directory exists. Anything that must not be read needs authentication.

It does not remove a page from search results. Disallow stops crawling, not indexing: a blocked URL that is linked from elsewhere can still appear, showing the URL with no snippet. To keep a page out of results, let it be crawled and serve <meta name="robots" content="noindex"> or an X-Robots-Tag header. Blocking it in robots.txt actually prevents the removal, because the crawler never sees the noindex.

It does not bind anyone. Compliance is voluntary. The crawlers of the large search and AI companies do honour it — that is why the list here uses their documented tokens — but a scraper that ignores it is not doing anything the file can prevent. Rate limiting and a WAF are the tools for that.

Blocking AI crawlers, accurately

The tokens matter more than the intent. A group for ChatGPT or OpenAI matches nothing at all; the names to use are GPTBot, OAI-SearchBot and ChatGPT-User, and they do different jobs — training, search indexing, and fetching a page because a user asked about it. Blocking all three is a different decision from blocking the first.

The same split runs through the list. Google-Extended controls Gemini training and has no effect on Search ranking or on whether you appear in AI Overviews; Googlebot is a separate matter entirely. Applebot-Extended governs Apple Intelligence training while Applebot keeps powering Siri and Spotlight. CCBot is Common Crawl, whose corpus is where most models started — blocking it now does not remove what was collected before.

Rules that are easy to get wrong

FAQ

Where does the file go?

At the root of every host and scheme you serve: https://example.com/robots.txt covers https://example.com and nothing else — not www.example.com, not the http:// version, not a subdomain. Each one needs its own file, and a 404 for it means "everything allowed".

Should I block my staging site here?

Belt and braces: yes, but the file is not what keeps it private. Put HTTP authentication in front of a staging host. A robots.txt on staging that gets copied to production with Disallow: / intact is one of the classic ways a site disappears from search — so check the deployed file after any release that touches it.

Does the order of the rules matter?

No, for Google, Bing and any crawler following the RFC: the longest matching rule wins and Allow breaks a tie of equal length. Some smaller crawlers still use first-match-wins, so if a bot's behaviour surprises you, write the rules so both interpretations agree.

Is anything sent anywhere?

No. The file is assembled in the page; nothing is fetched and nothing is uploaded. Test the result against a URL list in the robots.txt tester, which also runs locally.