The llms.txt Question Nobody Answers Properly
Everyone is shipping llms.txt files. Almost nobody is shipping one that changes whether a model can actually answer a question about them.
William Spurlock Founder — Spurlock Studios Updated 12 MIN
There is a specific failure mode I keep finding on sites that have “done AEO.” The llms.txt file exists. It returns 200. It is also completely useless, because it was written as a sitemap when it needed to be written as a briefing.
This sits under the Answer Engine Optimization playbook. The parent is the whole citation stack. This post is the file people treat as the whole stack. Crawler allow/deny still lives in AI crawlers and robots.txt. How Spurlock Studios scopes an on-site truth pass is on /visibility.
The short answer
- A crawler already has your sitemap.
llms.txthas to say what is true, who you are, and which URLs settle which questions. - Write an entity statement first: sector, specificity, location, one credential. Then point to pages that actually contain those facts.
- Every claim in the file should be checkable on your own domain. Marketing register is a token tax.
- Do not ship the file against client-rendered shells. If the linked page needs JavaScript to exist, you pointed at a locked door.
- The file is a living summary. Stale beats absent for damage, because it contradicts your other sources.
What is llms.txt actually for?
A crawler already has your sitemap. What it does not have is a compressed, unambiguous statement of what you do, who you are, and which pages settle which questions. That is the job. The format is a courtesy; the content is the product.
People keep confusing the jobs. A sitemap is an inventory of URLs a crawler is invited to fetch. robots.txt is a policy about who may fetch. JSON-LD is a graph of typed facts on the HTML page. llms.txt is a briefing: if a model only gets a few hundred tokens about you, what should those tokens be, and where should it go next to verify them?
If you paste the sitemap into /llms.txt, you have told the model that every URL is equally load-bearing. That is a lie. Your powder-coating spec page and your 2019 holiday recap are not the same kind of source. The briefing exists to rank the sources for a question, not to list the files.
The file does not rank you. It does not block training bots. It does not replace Organization schema. It does not make a thin About page suddenly citeable. It is a pointer with a caption. If the caption is vague and the pointer is a JavaScript shell, you shipped a 200 that changed nothing.
The longer spec-and-skip list is in llms.txt for brands. This post is the question I actually get: we added the file, why does it still not help?
What does a briefing look like compared to a sitemap dump?
Compare these two openings.
# Acme
- /about
- /services
- /contact
- /blog
# Acme Industrial Coatings
> Powder coating and industrial finishing for aerospace and defence
> subcontractors. AS9100D certified. Based in Wichita, KS. Founded 1998.
## What we do
- [Powder coating](/services/powder-coating): AS9100D certified line,
parts up to 4m, 48-hour turnaround.
- [Passivation](/services/passivation): Nitric and citric, per AMS 2700.
## Who we are
- [Dana Ruiz, founder](/team/dana-ruiz): 27 years in aerospace finishing.
The first tells a model where to look. The second tells it what is true. Only one of those survives being summarised into a single sentence by a system that is deciding whether to name you.
Read them the way a compressor reads them. The sitemap dump collapses to “Acme has an about page.” The briefing collapses to “Acme Industrial Coatings is an AS9100D powder coater in Wichita.” One of those sentences can be a citation. The other is a shrug.
The second example is also honest about scope. It does not claim to be every industrial finisher. It does not say “industry-leading.” It names a certification, a place, a year, a founder, and two service pages that are supposed to carry the same facts in HTML. That last part is the whole trick. The file is a table of contents for claims, not a place to invent claims that the site does not support.
If you cannot fill that second shape without lying, you do not have an llms.txt problem. You have an entity problem. Fix the pages.
Which three properties make the file useful?
Resolvable claims. Every line should be checkable against a page on your own domain. Models weight self-consistency heavily, and a claim in llms.txt that appears nowhere else reads as noise. If the file says 48-hour turnaround and the service page says “typical lead times vary,” you taught the model that your own sources disagree. Disagreement is how you get omitted, or worse, how you get named with a hedge.
Write the HTML first. Then summarise it. If you find yourself putting a fact in the briefing that you would not put on the page, delete it. The briefing is not a press release that crawlers are too polite to challenge.
Disambiguation up front. If your company name collides with anything — a band, a town, a bigger company in another sector — you resolve that in the first two lines or you lose the entity to whoever is more famous. Sector, location, and founding date do more work here than any adjective. “Acme” is a cartoon and a thousand LLCs. “Acme Industrial Coatings, Wichita, AS9100D, founded 1998” is a thing a model can keep separate from the crate.
This is the same work as entity architecture for AI search, just compressed. Preferred name, type, place, one credential. Repeat it until your About page, your Organization schema, and this file cannot be describing two companies.
No marketing register. “Industry-leading” is a token cost with no informational payload. Every word that a model cannot verify is a word competing with one it can. Superlatives get dropped. Specifics get kept. If you need the sentence to survive a summary, put a fact in it: a standard, a geography, a process, a named person, a page that proves it.
A useful-file checklist:
- Opening lines name sector, place, and one checkable credential
- Each bullet points at a URL that states the same fact in HTML
- Linked pages render without client-side JavaScript
- No adjective that would vanish if you asked “says who?”
- Name collision is resolved before the first service list
- Someone is assigned to update the file when the company changes
Where do teams get llms.txt wrong?
The most common mistake is treating the file as a one-time deliverable. Your llms.txt describes a company that changes. When you add a service line, the file is stale, and stale beats absent for damage because it introduces a contradiction between your own sources. Absent is a missing briefing. Stale is a briefing that argues with your homepage. Models are conservative when sources conflict. They hedge, or they pick the older, better-corroborated third party, or they skip you.
Put a calendar reminder on it. Better: make the file part of the same pull request that ships a new service page. If the HTML changed and the briefing did not, the PR is incomplete. That is a process rule, not a tooling problem.
The second most common mistake is shipping it without the markdown mirrors. If the file promises /services/passivation explains passivation, and that URL serves a JavaScript shell that resolves client-side, you have pointed a crawler at a locked door and told it there is a room behind it. Fetch the URL with JavaScript off. If the claim is not in the HTML, it is not in the briefing either, no matter what the file says.
Other ways to waste the 200:
- Dumping every blog URL so the “important pages” list is 400 lines of noise.
- Copying the meta description, which was already written for a search snippet, not for a model that needs disambiguation.
- Using the file to block bots. That is
robots.txt. Mixing policy and briefing means both jobs get done badly. - Inventing a founder bio that the /team page does not support. Self-inconsistency is louder than silence.
What is the right build order for llms.txt?
- Write the two-line entity statement. Sector, specificity, location, one credential.
- Ship JSON-LD for Organization and Person first —
llms.txtis a summary of a graph that needs to exist. - Confirm every linked page is server-rendered and reachable without JavaScript.
- Then write the file, and put a calendar reminder on it.
Doing it in that order takes a day longer and is the difference between a file that exists and a file that works.
Teams invert this because the file is fashionable and the graph is not. They add /llms.txt, screenshot the 200, and call the AEO ticket done. The model that actually answers a buyer question still has to pick a page. If the page is a shell, or the Organization node says a different city than the briefing, the file did not participate in the answer. It participated in the checklist.
Build order, said as a refusal list:
- Do not write the file before you can say the entity statement out loud without adjectives.
- Do not link a page you have not fetched with JS disabled.
- Do not add a service bullet before the service URL states the same facts.
- Do not treat a 200 as the acceptance test. The acceptance test is: a stranger can recover who you are and where to verify it from the file alone, and then confirm it on the HTML.
That is the whole properly. Not a new format. The same facts, in the right order, kept in sync.
Is llms.txt just a sitemap in markdown?
No. A sitemap inventories URLs. llms.txt briefs a model on which of those URLs settle which questions, and what is true before it fetches anything. If your file is a bullet list of paths with no captions and no entity statement, you shipped a sitemap with a fashionable filename. The crawler already had that list.
Should you ship llms.txt before JSON-LD?
No. The file is a summary of a graph that needs to exist. Organization and Person on the HTML, in JSON-LD, with the same names and places you will put in the briefing, come first. Then the file. A briefing that points at an empty graph is a caption with no picture, and models are not obligated to trust captions.
What questions does this article answer?
- Is llms.txt just a sitemap in markdown?
- No. A sitemap inventories URLs. `llms.txt` briefs a model on which of those URLs settle which questions, and what is true before it fetches anything. If your file is a bullet list of paths with no captions and no entity statement, you shipped a sitemap with a fashionable filename. The crawler already had that list.
- Should you ship llms.txt before JSON-LD?
- No. The file is a summary of a graph that needs to exist. Organization and Person on the HTML, in JSON-LD, with the same names and places you will put in the briefing, come first. Then the file. A briefing that points at an empty graph is a caption with no picture, and models are not obligated to trust captions.
Last reviewed
AI Visibility
AI Visibility When ChatGPT names the franchise, not your shop
Run the best-HVAC-near-me prompt panel. If the model names a national franchise, fix corroboration and entity facts — not another blog calendar.
AI Visibility How do I get cited by Perplexity specifically
Allow PerplexityBot, put a liftable answer and unique numbers in HTML, then log numbered sources on a frozen prompt panel. There is no bought citation rate.
AI Visibility What belongs in an AI visibility monthly retainer vs a one-time audit
A one-time audit is the baseline plus prioritized fixes. A monthly retainer is prompt-panel tracking, entity hygiene, page jobs, and citation recovery.
AI Visibility Does Wikipedia or Wikidata help AI recommend my brand
Wikipedia is not a paid AI lever. Notability plus independent sources decide the page; a real Wikidata item helps entity consistency, not a promotional stub.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.