Skip to content

Guide

llms.txt explained: what it is and how to publish it

llms.txt is a plain Markdown file at the root of a website that gives a language model a short, curated map of the site. This guide covers the format, a worked example, how to publish one, and an honest account of what is and is not known about who reads it.

By the Toolfound team. Last reviewed

What llms.txt is

The idea was proposed in September 2024 by Jeremy Howard and is described at llmstxt.org. The problem it addresses is practical: a whole website rarely fits in a model's context window, and ordinary HTML is padded with navigation, scripts and layout. A curated Markdown index lets a reader, human or machine, start from the pages that matter.

It is a community proposal, not an Internet standard. No standards body owns it, and a site owner is free to publish one or not.

How llms.txt differs from the files it is compared to
FileWritten forWhat it says
robots.txtCrawlersWhich paths a crawler may or may not request.
sitemap.xmlSearch enginesEvery URL you want discovered, with dates.
llms.txtLanguage models and the people configuring themThe few pages worth reading first, each with a note on what it answers.

The three do different jobs and none replaces another. Publishing llms.txt does not allow or block anything.

The format, in order

The file is Markdown, and the specification defines its parts in a fixed order:

  1. An H1 with the name of the site or project. This is the only required element.
  2. A blockquote with a short summary of the site, enough to understand the rest of the file.
  3. Zero or more paragraphs or lists with more detail. These must not contain headings.
  4. Zero or more H2 sections, each a Markdown list of links. Each item is [name](url), optionally followed by a colon and a note about the page.
  5. A section titled Optional has a defined meaning: its links can be skipped when a shorter context is needed.

The specification also suggests that pages likely to be useful to a model offer a clean Markdown version at the same URL with .md appended. That part is a suggestion, and the file works without it.

A worked example

Here is a complete file for a made-up note-taking product. It uses every element of the format.

https://example.com/llms.txt
# Acme Notes

> Acme Notes is a note-taking app for small teams. It syncs across devices, supports Markdown, and costs $5 per user per month.

The app is hosted at app.example.com. Documentation lives under /docs. Notes are private by default.

## Product

- [Pricing](https://example.com/pricing): plans, limits and what each plan includes
- [Features](https://example.com/features): sync, Markdown, sharing, offline mode

## Docs

- [Quick start](https://example.com/docs/quick-start): create a workspace and invite a teammate
- [API reference](https://example.com/docs/api): endpoints, authentication and rate limits
- [MCP server](https://example.com/docs/mcp): connect an assistant to a workspace

## Policies

- [Privacy](https://example.com/privacy): what is stored and for how long

## Optional

- [Changelog](https://example.com/changelog): release notes by version

The summary says what the product is, who it is for and what it costs, with no adjectives. Every link has a note that names the question the page answers, so a reader can choose without opening it. URLs are absolute because the file is read away from the site that serves it. The changelog sits under Optional because it is useful but not needed to understand the product.

You can build a file like this with the llms.txt generator, or paste an existing one into its validator to check the structure.

What AI assistants actually do with it

State of knowledge on 2026-10-10: llms.txt is a proposal, and we know of no search engine or AI assistant provider that has published a commitment to read it for ranking or citation. Some developer tools and agents do read it when a user points them at a site, because a short index is cheap to load. Beyond that, claims that it makes assistants cite you are not established.

Toolfound serves an llms.txt, and in September 2026 we also counted a large number of ChatGPT fetches of our pages. Our measurements do not show whether any of those visits was prompted by the file, so we do not claim it was.

What can and cannot be said
StatementStatus
The format is publicly specified at llmstxt.org.True
A well-formed file costs a visitor or a crawler almost nothing to read.True
Major search engines use llms.txt as a ranking signal.No published commitment that we know of
Publishing one causes assistants to cite your site.Not established
Some tools read it when pointed at a site.True, tool by tool

The sensible position is to treat it as a low-cost, optional convenience. Write it because a short, honest description of your product is worth having in one place, and do not spend a week on it.

What to put in it

  • A summary that states what the product is, who it is for and the price, in plain words.
  • The pages a stranger needs first: pricing, features, quick start, API or integration docs, support and policies.
  • A note on every link that says what the page answers, not what it is called.
  • The address of your MCP server or API, if you have one, so an agent has a short route from the product to the interface.
  • Secondary material under an Optional section.

Leave out marketing superlatives, pages that need a login, and anything you would not want quoted back. A reader may take the file at its word.

How to publish it

  1. Write the file with the generator or by hand, and name it llms.txt.
  2. Serve it at the root of the domain, so https://yourdomain.com/llms.txt returns it with a 200 status. Plain text or Markdown in UTF-8 are both fine.
  3. On a static host, put the file in the public folder. On a framework, add a route that returns the text.
  4. Make sure robots.txt does not block the path and that the response is not your HTML app shell.
  5. Request it with curl -i https://yourdomain.com/llms.txt to confirm status and content type, then paste the body into the validator.
Next.js App Router: app/llms.txt/route.ts
export function GET() {
  const body = [
    "# Acme Notes",
    "",
    "> Acme Notes is a note-taking app for small teams.",
    "",
    "## Docs",
    "",
    "- [Quick start](https://example.com/docs/quick-start): create a workspace",
  ].join("\n");
  return new Response(body, {
    headers: { "Content-Type": "text/plain; charset=utf-8" },
  });
}

Update the file when your pages change. A link that now redirects or returns an error is worse than no link, and the redirect checker will show you which ones moved. If you run an MCP server, the MCP listing preflight also looks for this file on your site.

For the wider picture of being found by assistants, read get your software cited by AI assistants.

Common mistakes

  • Relative URLs such as /docs. The file is read out of context, so use full addresses.
  • Serving the site's HTML for a missing file. A 200 response full of markup is not an llms.txt.
  • Listing hundreds of URLs. A sitemap already does that; this file is a curated shortlist.
  • Headings below H2, or a second H1, which the format does not define.
  • Marketing slogans in the summary in place of a plain statement of what the product does.
  • Forgetting it exists. Stale links erode the point of a curated list.

Common questions

Is llms.txt an official standard?

No. It is a community proposal published at llmstxt.org. No standards body maintains it, and publishing one is optional.

Does llms.txt replace robots.txt or my sitemap?

No. robots.txt controls which paths crawlers may request and the sitemap lists URLs for search engines. llms.txt is a short, annotated shortlist of the pages worth reading first. It neither allows nor blocks anything.

Will publishing one get my site cited by AI assistants?

That is not established. We know of no major provider that has committed to using it for citation or ranking. It is cheap to publish, so it is a reasonable convenience, but it is not a lever to count on.

Where does the file go and what should it be called?

At the root of the domain as llms.txt, so it resolves at https://yourdomain.com/llms.txt, served with a 200 status as plain text or Markdown.

How long should the file be?

Short. One summary, a handful of sections and the few links a newcomer needs. Past a few dozen links it stops being an index.

Do I need the .md versions of every page?

No. The specification suggests offering clean Markdown versions of useful pages at the same URL with .md appended, but the llms.txt file works on its own.