Key takeaways
- llms.txt is a single markdown file at your domain root that lists the pages you most want AI answer engines to read, in a fixed order: H1 name, blockquote summary, then link sections
- It is a proposal from llmstxt.org, not a ratified standard, and no major AI provider has publicly committed to reading it, so treat it as a cheap bet rather than a ranking factor
- It does not control access (that is robots.txt's job) and it cannot make a model cite you
- The file must be genuinely curated: twenty links a human would actually pick beats a dump of your sitemap, because the whole point is a shorter, higher-signal context
llms.txt is a proposed standard: a single markdown file at your domain root that hands AI models a short, curated map of your most useful pages. It is a genuinely interesting idea, it costs about twenty minutes to implement, and it is also widely misexplained, usually with more confidence than the evidence supports. This guide covers what the file actually is, what it demonstrably is not, how to write a good one, and how to decide whether it deserves a slot on your roadmap.
What llms.txt is
llms.txt is a proposed convention, published at llmstxt.org, for a single markdown file at the root of your domain (https://example.com/llms.txt) that gives a large language model a short, curated map of your most useful content.
The reasoning behind it is straightforward. When a model answers a question about your site, it works from a limited context window. A full crawl of a 4,000-page site is far too much material, and the model has no reliable way to tell your flagship documentation from a tag archive. llms.txt is your chance to answer the question "if you could only read fifteen pages of this site, which fifteen?"
The format is deliberately minimal and strictly ordered:
# Example Ltd
> Example builds accounting software for UK sole traders and micro-businesses.
Prices are in GBP and include VAT. Support hours are 09:00-17:00 UK time.
## Docs
- [Getting started](https://example.com/docs/start): Set up an account and file your first return.
- [VAT returns](https://example.com/docs/vat): How quarterly VAT submission works, end to end.
## Optional
- [Changelog](https://example.com/changelog): Release notes, updated weekly.
Four parts, and the order matters:
- An H1 with the site or project name. Required, and there must be exactly one.
- A blockquote giving a one-sentence summary. Optional but strongly recommended; it is the single highest-value line in the file.
- Free-form markdown for anything a model should know before reading the links: currency, jurisdiction, licensing, naming conventions.
- H2 sections containing link lists. Each item is a markdown link followed by an optional
: description.
The one section name with defined meaning is ## Optional. Links under it are explicitly marked as safe to skip when the model needs a shorter context. Everything else is yours to name.
Tip
The blockquote is worth more thought than the link list. It is the line most likely to be quoted back at a user who asks "what is Example Ltd?", so write it as a factual, self-contained sentence, not as marketing copy.
What llms.txt is not
This is where most of the confusion lives, and it is worth being blunt about each one.
It is not an access-control file. llms.txt grants nothing and blocks nothing. If you want to stop GPTBot from crawling you, that is a Disallow in robots.txt; publishing an llms.txt has no bearing on it. The two files are complements, not alternatives.
It is not a ratified standard. It is a proposal, first published in 2024, adopted voluntarily. As of August 2026 no major AI provider has publicly documented that its crawlers read llms.txt. Anyone telling you it is "required" is guessing.
It is not a sitemap. sitemap.xml is exhaustive by design: every indexable URL, so a search engine can discover all of them (an XML sitemap checker is the tool for auditing that one). llms.txt is the opposite: it is valuable precisely because it is short. Generating it from your sitemap destroys the only thing it was for.
It is not a ranking factor. There is no published evidence that llms.txt affects Google, Bing, or any AI engine's rankings or citation behaviour. Treat any confident claim otherwise as unsupported.
Heads up
Be careful with the "AI SEO" pitch that llms.txt will get you cited in ChatGPT. Nobody has demonstrated that. The honest case for the file is cheaper: it costs very little, it forces a useful editorial exercise, and if adoption does arrive you already have one.
So is it worth publishing?
For most sites, yes, but for unglamorous reasons.
The cost is close to zero: one static file, no build step, no maintenance beyond an occasional review. The upside, if adoption comes, is that you have already shaped how models describe you. And there is a real side benefit that has nothing to do with AI at all: being forced to name the fifteen pages that matter is a genuinely clarifying exercise. Teams that do it usually discover the list is harder to write than expected, which is itself a finding about the site.
The case against is honest too. If your roadmap has actual technical debt on it (broken canonicals, redirect chains, pages that return the wrong status code), those affect real crawlers today, and llms.txt affects hypothetical ones tomorrow. Fix the certain problem first.
How to write a good llms.txt
Curate ruthlessly. Ten to thirty links is the useful range. If you are over fifty, you have written a sitemap.
Describe every link. The : description after each link is where a model learns what it is about to read. - [VAT returns](/docs/vat) tells it almost nothing; - [VAT returns](/docs/vat): How quarterly VAT submission works, end to end tells it a lot.
Use absolute URLs. Relative paths are ambiguous once the file is pulled out of its original context.
Group by intent, not by your CMS. ## Docs, ## Pricing, ## Policies are useful section names. ## Posts and ## Pages are not.
Put nice-to-haves under ## Optional. This is the only section name with defined behaviour, and it lets a model degrade gracefully when the context is tight.
Serve it correctly. Plain text or markdown, HTTP 200, no redirect, no login wall, no JavaScript rendering required. It is a static file; treat it like one.
Write facts, not adjectives. A model reading "the leading platform for X" gets no usable information. "Accounting software for UK sole traders, from £12/month" is something it can actually repeat.
If you want a valid file without hand-writing the syntax, Crawl Cove's llms.txt generator builds one in the browser: add your name, summary and sections, then copy or download the result. It runs entirely in your browser, so nothing is uploaded anywhere.
Checking it after you publish
Two things are worth verifying, and neither is automatic.
The first is simply that the file is reachable: https://yourdomain.com/llms.txt should return a 200 with a text content type, from a fresh request rather than your browser cache.
The second is that the links inside it are still good. An llms.txt full of URLs that now 404 or redirect is worse than none, because it hands a model a curated list of broken pages. This is the failure mode to watch for: the file is static, but the site it describes is not, and nothing warns you when they drift apart. Our free AI crawler access checker is a quick way to spot-check a single URL's bot access while you're at it, before you widen the check to the whole site. For the wider single-page picture, the free AI search visibility checker reads the same URL and reports whether /llms.txt exists alongside crawler access, snippet controls and structured data.
Crawl Cove flags a missing /llms.txt as a low-severity advisory in its AI section during a normal site audit, alongside its check for whether GPTBot, ClaudeBot, PerplexityBot or Google-Extended are disallowed in your robots.txt. That section is what our AI search visibility checker reports on: the file itself is only one signal, and it sits next to whether your pages are structured so a model can extract an answer from them at all. Because it crawls every URL on the site in the same pass, the links you listed get their status codes and redirect paths checked as part of the same audit rather than as a separate chore. It runs locally on your own machine, so you can re-run it after each edit and compare audits to confirm the file and the site still agree. You can try it free for 14 days.
Our guide to measuring AI visibility covers both halves of that page: these deterministic readiness checks and the manual AI-citation tracker alongside them.
Related reading
- AI crawlers and robots.txt: who to allow and who to block covers the access side of the same question.
- Technical SEO: the complete guide shows where files like this sit in the wider picture.
- Crawl budget explained explains why "which pages matter most" is a question conventional crawlers ask too.
Wrap-up
llms.txt is a small, cheap, sensible idea with an uncertain future. It is a curated markdown map of your best pages, it lives at your domain root, and it is not an access control, a sitemap, or a ranking factor. Publish one if you can spare twenty minutes: write a factual summary, list the pages a knowledgeable colleague would list, describe each one in a clause, and keep the links alive. If adoption never arrives you have lost almost nothing. If it does, you have already answered the question every model will be asking.
Frequently asked questions
- Is llms.txt an official standard?
- No. It is a community proposal published at llmstxt.org in 2024 and adopted voluntarily by a growing number of sites. No major AI provider has publicly confirmed that it reads the file, so publish it as a low-cost bet, not as a ranking tactic.
- Does llms.txt replace robots.txt or sitemap.xml?
- No, and it is not a substitute for either. robots.txt states which crawlers may fetch which URLs, sitemap.xml lists every indexable URL for search engines, and llms.txt curates the handful of pages you most want a model to read. They answer three different questions.
- Where does the llms.txt file go?
- At the root of your domain, served at https://yourdomain.com/llms.txt as plain text or markdown with a 200 response. Same convention as robots.txt.
- Will publishing llms.txt improve my Google rankings?
- There is no evidence that it will, and you should be sceptical of anyone claiming otherwise. Google has not said it reads llms.txt, and Google-Extended, its AI opt-out token, lives in robots.txt, not here.