urltoolskit.org
URL utilities, in the browser
Say hi →

llms.txt Generator

/llms.txt · markdown · built and validated as you type

robots.txt tells a crawler what it may not fetch. llms.txt tells a language model what is worth fetching: your title, one line about what this site is, and a curated list of pages with a sentence each. It is markdown with two rules that are easy to break — exactly one H1 at the top, and no heading deeper than H2 — so this builds it from whatever shape your URL list is in, and checks it.

What llms.txt is

A markdown file at the root of your site — https://example.com/llms.txt — proposed in 2024 as a way to hand a language model a short, curated map of a site instead of leaving it to crawl the navigation, the cookie banner and the footer. The format is deliberately tiny:

# Project name

> One-sentence summary of what this is.

Any amount of ordinary markdown, with no headings in it.

## Docs

- [Quickstart](https://example.com/docs/quickstart): five minutes to a first request
- [API reference](https://example.com/docs/api)

## Optional

- [Blog](https://example.com/blog)

The H1 is the only required element. Everything else is optional, and the sections are just H2 headings with a bullet list of markdown links under them, each with an optional description after a colon.

The two rules people break

Paste your links in whatever shape you have them

The links box accepts a mix, one entry per line, and works out which is which:

Lines with no URL in them are skipped and reported, rather than being emitted as broken bullets.

The "Optional" section

## Optional is the one section name with a defined meaning: links under it may be skipped when a shorter context is needed. Everything else is a name you choose. If you have several sections and none of them is Optional, this says so — deciding what a model can afford to skip is the part of the file that does the work.

What gets checked

Publishing it

  1. Serve the file at /llms.txt — the root, not /.well-known/, which is where security.txt goes.
  2. Send it as text/markdown; charset=utf-8, or text/plain if your host will not do the former. Both are read fine; neither should be text/html.
  3. Keep it small. It is an index, not a corpus: a link with a good sentence beats ten links without one.
  4. The related llms-full.txt convention puts the actual content of those pages in one file. That is a different job — this tool builds the index.

FAQ

Does anything actually read it?

Adoption is real but partial: it is a community proposal, not a standard, and no major crawler has committed to honouring it the way they honour robots.txt. It costs one small file, it is human-readable, and it is the only artefact that currently expresses "here is what my site is for". Treat it as cheap insurance rather than a guarantee.

Is it a replacement for robots.txt or a sitemap?

No. robots.txt is permission, a sitemap is completeness, and llms.txt is curation — the twenty pages that matter, described. They answer different questions and a site can have all three.

Can I block AI crawlers and still publish this?

Yes, and it is a coherent position: block the training crawlers in robots.txt, and keep an llms.txt for the assistants a user has pointed at your docs on purpose. The robots.txt generator has one-click blocks for the named AI user-agents.

Is anything uploaded?

No. The file is built, validated and downloaded entirely in your browser. Nothing is fetched and no URL you paste is visited — which also means link rot is not checked here.