llms.txt Generator
robots.txt tells a crawler what it may not fetch. llms.txt tells a language model what is worth fetching: your title, one line about what this site is, and a curated list of pages with a sentence each. It is markdown with two rules that are easy to break — exactly one H1 at the top, and no heading deeper than H2 — so this builds it from whatever shape your URL list is in, and checks it.
What llms.txt is
A markdown file at the root of your site — https://example.com/llms.txt — proposed in 2024 as a way to hand a language model a short, curated map of a site instead of leaving it to crawl the navigation, the cookie banner and the footer. The format is deliberately tiny:
# Project name
> One-sentence summary of what this is.
Any amount of ordinary markdown, with no headings in it.
## Docs
- [Quickstart](https://example.com/docs/quickstart): five minutes to a first request
- [API reference](https://example.com/docs/api)
## Optional
- [Blog](https://example.com/blog)
The H1 is the only required element. Everything else is optional, and the sections are just H2 headings with a bullet list of markdown links under them, each with an optional description after a colon.
The two rules people break
- Exactly one H1, and it comes first. A file that opens with a comment, a badge line or a table of contents fails the first thing a parser checks.
- No heading below H2. An H3 inside a section does not create a sub-section — for a consumer walking H2 to H2, it is unstructured text in the middle of a link list. This flags it as a problem rather than a style note.
- Links must be absolute. The file is read away from your server, often by something that never saw the page it came from.
/docs/apiresolves to nothing there. Give the site URL above and relative links are expanded for you.
Paste your links in whatever shape you have them
The links box accepts a mix, one entry per line, and works out which is which:
## Section name— starts a new section. Everything after it goes under that heading.- [Title](https://url): description— already markdown, used as-is.https://example.com/docs/api: the reference— the description follows the colon.API reference — https://example.com/docs/api— title first, separated by a dash, comma or pipe.https://example.com/docs/getting-started— no title at all, so one is made from the last path segment: Getting started.
Lines with no URL in them are skipped and reported, rather than being emitted as broken bullets.
The "Optional" section
## Optional is the one section name with a defined meaning: links under it may be skipped when a shorter context is needed. Everything else is a name you choose. If you have several sections and none of them is Optional, this says so — deciding what a model can afford to skip is the part of the file that does the work.
What gets checked
- A single H1, first in the file; no H3 or deeper anywhere.
- Every bullet is a real markdown link with title text.
- Every URL is absolute;
http://is flagged, and can be rewritten tohttps://. - Duplicate links across all sections — dropped or reported, your choice.
- Empty sections, a missing summary, and a details block that smuggled in a heading.
- Tracking parameters (
utm_*,fbclid,gclidand the rest) stripped, because a curated map should not be carrying campaign tags.
Publishing it
- Serve the file at
/llms.txt— the root, not/.well-known/, which is where security.txt goes. - Send it as
text/markdown; charset=utf-8, ortext/plainif your host will not do the former. Both are read fine; neither should betext/html. - Keep it small. It is an index, not a corpus: a link with a good sentence beats ten links without one.
- The related
llms-full.txtconvention puts the actual content of those pages in one file. That is a different job — this tool builds the index.
FAQ
Does anything actually read it?
Adoption is real but partial: it is a community proposal, not a standard, and no major crawler has committed to honouring it the way they honour robots.txt. It costs one small file, it is human-readable, and it is the only artefact that currently expresses "here is what my site is for". Treat it as cheap insurance rather than a guarantee.
Is it a replacement for robots.txt or a sitemap?
No. robots.txt is permission, a sitemap is completeness, and llms.txt is curation — the twenty pages that matter, described. They answer different questions and a site can have all three.
Can I block AI crawlers and still publish this?
Yes, and it is a coherent position: block the training crawlers in robots.txt, and keep an llms.txt for the assistants a user has pointed at your docs on purpose. The robots.txt generator has one-click blocks for the named AI user-agents.
Is anything uploaded?
No. The file is built, validated and downloaded entirely in your browser. Nothing is fetched and no URL you paste is visited — which also means link rot is not checked here.