robots.txt generator
A robots.txt is four directives and a great many ways to get them subtly wrong. Fill in the paths you want kept out of the crawl, tick the AI crawlers you would rather not feed, and take the file away — checked, as it is written, by the same rules the tester applies.
Block AI crawlers
How to use
- Pick a preset if one matches, then edit it. The presets are starting points, not opinions — a shop that has no facets does not need the facet rules.
- Put one path per line under Disallow. A full URL pasted in is reduced to its path, and a missing leading slash is added, because both are the usual way this file gets broken.
- Use Allow for exceptions inside a disallowed path. Ordering does not matter: the longest matching rule wins and Allow breaks ties, so
Allow: /admin/public/beatsDisallow: /admin/. - Tick the AI crawlers you want to exclude. Each gets its own group, because a bot-specific group replaces the
*group entirely rather than adding to it. - Add your Sitemap URLs — absolute, with the scheme. This is the one line in the file that helps rather than restricts.
- Test a path against what you have built, then copy or download and put the file at
https://yourdomain/robots.txt. Nowhere else: a subdirectory robots.txt is never read.
What a robots.txt does not do
It is not access control. The file is public, and listing /internal-reports/
in it tells everyone the directory exists. Anything that must not be read needs authentication.
It does not remove a page from search results. Disallow stops crawling, not
indexing: a blocked URL that is linked from elsewhere can still appear, showing the URL with no
snippet. To keep a page out of results, let it be crawled and serve
<meta name="robots" content="noindex"> or an X-Robots-Tag header.
Blocking it in robots.txt actually prevents the removal, because the crawler never sees
the noindex.
It does not bind anyone. Compliance is voluntary. The crawlers of the large search and AI companies do honour it — that is why the list here uses their documented tokens — but a scraper that ignores it is not doing anything the file can prevent. Rate limiting and a WAF are the tools for that.
Blocking AI crawlers, accurately
The tokens matter more than the intent. A group for ChatGPT or OpenAI
matches nothing at all; the names to use are GPTBot, OAI-SearchBot and
ChatGPT-User, and they do different jobs — training, search indexing, and fetching a
page because a user asked about it. Blocking all three is a different decision from blocking the
first.
The same split runs through the list. Google-Extended controls Gemini training and has
no effect on Search ranking or on whether you appear in AI Overviews; Googlebot is a
separate matter entirely. Applebot-Extended governs Apple Intelligence training while
Applebot keeps powering Siri and Spotlight. CCBot is Common Crawl, whose
corpus is where most models started — blocking it now does not remove what was collected before.
Rules that are easy to get wrong
Disallow: *means two different things. Google's matcher reads the bare*as a wildcard from the start of the path, so it blocks the entire site; a crawler that requires a leading/ignores the line and crawls everything. WriteDisallow: /when you mean the whole site, and a real path when you do not.- Wildcards work in paths, not in user-agents.
Disallow: /*?sort=is a wildcard;User-agent: Google*is a plain substring and matches nothing extra. $anchors, but only at the end.Disallow: /*.pdf$blocks PDFs; a$anywhere else is a literal dollar sign.- A group ends at the first rule after a run of user-agent lines. Several
User-agentlines in a row share the rules beneath them; a blank line between them splits the group. - Case matters in paths and not in directives.
disallow:is fine;/Admin/and/admin/are different URLs. - Crawl-delay is ignored by Google. Bing and Yandex read it. For Google, set the crawl rate in Search Console.
FAQ
Where does the file go?
At the root of every host and scheme you serve: https://example.com/robots.txt covers https://example.com and nothing else — not www.example.com, not the http:// version, not a subdomain. Each one needs its own file, and a 404 for it means "everything allowed".
Should I block my staging site here?
Belt and braces: yes, but the file is not what keeps it private. Put HTTP authentication in front of a staging host. A robots.txt on staging that gets copied to production with Disallow: / intact is one of the classic ways a site disappears from search — so check the deployed file after any release that touches it.
Does the order of the rules matter?
No, for Google, Bing and any crawler following the RFC: the longest matching rule wins and Allow breaks a tie of equal length. Some smaller crawlers still use first-match-wins, so if a bot's behaviour surprises you, write the rules so both interpretations agree.
Is anything sent anywhere?
No. The file is assembled in the page; nothing is fetched and nothing is uploaded. Test the result against a URL list in the robots.txt tester, which also runs locally.