Spurlock Studios
Contact
Share LinkedIn X
A brushed metal coupon. Thesis: LLMS TXT BRANDS PUBLISH SKIP.

llms.txt should brief a model on who you are and which URLs settle which questions. It is not a sitemap dump, not a ranking hack, and not a file every answer engine promises to read. Jeremy Howard’s llmstxt.org proposal (September 2024; v2 on the same site) is a markdown convention for inference-time context — a compact packet agents can fetch, then follow links from.

This spoke sits under the Answer Engine Optimization playbook. Use it when you are ready to ship or rewrite the file as part of your on-site truth layer. Crawler allow/deny still lives in AI crawlers and robots.txt.

The short answer

  • Publish /llms.txt as a briefing: legal name, offer, ICP or geography, and a short list of pages that settle buyer questions.
  • Follow the spec order: H1 (the only required section), optional blockquote summary, then H2 “file lists” of [name](url): notes.
  • Do not paste the sitemap. sitemaps.org lists indexable URLs for crawlers. llms.txt curates a handful of answer pages for agents.
  • Adoption is uneven. Chrome’s Lighthouse llms.txt audit calls the file an “emerging convention” and marks a 404 as N/A because the file is optional.
  • Pair it with robots.txt (RFC 9309). Training bots and search bots are different agents — see OpenAI’s crawler overview and Anthropic’s crawler help article.

What is llms.txt — and what is it not?

The convention is simple: publish /llms.txt (and optionally /llms-full.txt or a path-scoped /docs/llms.txt) in markdown so tooling can fetch a human-authored summary. Howard’s original write-up is the Answer.AI post of 3 September 2024. The living format lives at llmstxt.org. The repo is AnswerDotAI/llms-txt.

It is not an IETF or W3C standard. It is not a robots exclusion file. It is not a promise that ChatGPT, Claude, Gemini, or any other product will load the file on every brand prompt.

JobFileWho it is forWhat it is not
Brief identity + point to answer URLs/llms.txtAgents and tools that fetch a summaryA ranking lever or training opt-out
List indexable URLs/sitemap.xmlSearch crawlersA curated briefing
Allow or deny crawl / product tokens/robots.txtCompliant crawlersA place to explain who you are
Machine-readable entity factsHTML + schema.org OrganizationSearch and extractorsA substitute for the briefing

If your only AEO move is uploading a thin llms.txt, you will be disappointed. If you use it to encode the same facts you want models to repeat — and you put those facts on the linked HTML pages — it earns its keep even when a given session never fetches the file.

Decision list:

  1. Need a compact “who we are / start here” packet? → write llms.txt.
  2. Need every public URL discovered? → sitemap.
  3. Need to block GPTBot or allow OAI-SearchBot? → robots.txt, not this file.

What does the llms.txt spec actually require?

llmstxt.org is picky about order, not about brand poetry. A file at /llms.txt (or a subpath such as /docs/llms.txt) covers URLs under that path. Where more than one file applies, the spec says agents should use the most specific one.

Required and optional sections, in this order:

OrderSectionRequired?Brand use
1Optional BOMNoSkip unless your CMS injects one
2# H1 — project or site nameYes — only required sectionLegal or primary trade name
3> blockquote summaryNo, but do itICP, geography, one-line offer
4Body copy with no headingsNoDisambiguation, “we are not X”
5## H2 file listsNoLinked answer pages with notes
6## OptionalConventionSecondary links agents can skip

Each file-list item is a markdown hyperlink, then optionally a colon and notes:

- [Outbound sequencing](https://example.com/services/outbound): Human-approved
  AI sequences for B2B teams with 5–50 SDRs; HubSpot and Salesforce.

The spec also proposes clean markdown twins of HTML pages (.md or .html.md) and rel="describedby" / rel="alternate" hints. Brands can ignore the twin-page machinery on day one. Do not ignore the H1-then-lists shape. A YAML dump, a JSON-LD blob, or a pasted sitemap.xml is not an llms.txt.

Howard considered RFC 8615 /.well-known/ and rejected it as the only home: many authors control a path, not the origin root. Path-scoped files are allowed. One file per registrable brand domain is still the sane default for a company site.

Why does an llms.txt briefing beat a sitemap dump?

sitemaps.org exists so crawlers can find URLs. A production sitemap often includes pagination, tag archives, thin landing pages, and last year’s campaign leftovers. That list is the wrong context window for “what does this company sell?”

The spec says the same thing in plainer language: a sitemap will generally cover documents that in aggregate will not fit in a model context, and it will include a lot of information that is not necessary to understand the site. It also will not list the external proof URLs you want an agent to read (a certification registry, a public filing, a partner directory).

Sitemap dumpBriefing
Every blog URL from five yearsOne method page that settles “how we work”
/blog/page/2Nothing — pagination is not an answer
Nav labels with no factsOne clause of scope per offer
400 URLs, no priority8–20 URLs, each tied to a question
Changes when you publish a postChanges when an offer, HQ, or fact changes

Checklist — if you are about to paste the sitemap, stop:

  • Can you name the question each URL answers?
  • Would you hand this list to a tired analyst with 30 seconds?
  • Does every link still 200?
  • Is the claim on the destination page, not only in the bullet?

If any box is unchecked, you are writing a second sitemap. Delete and start from questions.

Which URLs should llms.txt use to settle which questions?

This is the whole job. Pick the questions a buyer or a model will actually ask, then point at the one page that is allowed to answer.

Question a model will get askedURL that should settle itWhat the bullet must say
Who are you / what is the legal name?About or entity pageCanonical name, AKA, founding year if public
What do you sell?Service or product hub or one page per offerScope in one clause, not “solutions”
Who is this for / not for?ICP, industries, or “not a fit” sectionGeography, company size, exclusions
How do we start?Contact, audit, or booking pageOne commercial door
What does it cost?Pricing or packages — only if you will stand behind the number in chatRange or “custom, starts at”
Can I trust this?Case study, credentials, or compliance pageNamed proof, not “trusted by leaders”

Ten rows is plenty. Twenty is a stretch. If you cannot fill the middle column, the page does not exist yet — write the HTML first.

Procedure:

  1. List ten buyer questions in a spreadsheet. No marketing adjectives.
  2. Assign one production URL per question. Duplicates mean you have two pages fighting.
  3. Open each URL. If the fact is missing, add it to the page or drop the question.
  4. Rewrite survivors as spec-shaped bullets.
  5. Paste the draft into a blank chat with only the markdown. Ask: what do they sell, who should not hire them, which URL first?

If the model invents a fourth product line, your briefing leaked ambiguity. Fix the file before it hits production.

What should I put in llms.txt?

Include facts a model should be allowed to repeat without inventing:

  • Canonical brand name and any public “also known as”
  • Primary offers with plain-language scope
  • Service area or ICP boundaries
  • Links to pages that expand each claim
  • Notable, verifiable credentials (certifications, years, named partnerships)
  • Preferred commercial entry point
  • One exclusion line when wrong-fit recommendations are expensive

Keep each bullet one idea. Link the URL that settles the claim. If the claim is not on the destination page, either add it there or delete the bullet.

PublishExample note after the colon
Legal name + AKA“Trade name Northline; legal Northline Field Services LLC”
Offer + scope“Quarterly PM for rooftop units; SLA-backed response”
Geography“Southeast multi-site retail; not residential”
Credential“EPA Section 608 certified techs; COIs on /compliance”
Door“Request a multi-location assessment”

Verified studio receipts you may mirror in your file when they are also on HTML: 500+ automations built, 20,000+ hours on agentic systems, 35,000+ hours saved for clients, hundreds of production sites, Make.com AI automation certifications, SEO certified since 2021. Do not invent a parallel set of numbers for a client brand.

What should I skip in llms.txt?

Skip anything that teaches a model the wrong shape of the company.

SkipWhy it fails
Every blog URL from the last five yearsContext landfill; no question attached
Login, cart, and utility routesNot answers
Thin tag archives and parameter URLsDuplicate, low-signal
Superlatives with no proofCompresses into hallucinated mush
Pricing you will not stand behind in chatModels will quote it
Internal codenames and unreleased productsBecomes “they ship X” overnight
Competitor attack linesKeep the file factual; comparisons exist elsewhere
Contradictory draftsLinkedIn 2019 vs site 2017 — fix sources first

Also skip using llms.txt as a hiding place. If the page is public on LinkedIn and in ads, omitting it from the briefing does not unsay it. If you do not want a path retrieved, that is a robots.txt decision, not a briefing omission.

What does a brand llms.txt skeleton look like?

Documentation sites (FastHTML is the spec’s own example) list API pages. A brand file should read like a one-page analyst brief.

# Northline Field Services
> Commercial HVAC maintenance and retrofit for multi-site retailers
> across the Southeast. Founded 2009. EPA Section 608 certified techs.

Northline is not a residential HVAC franchise and is not affiliated
with Northline Robotics (Delaware).

## What we do
- [Planned maintenance](https://example.com/services/maintenance): quarterly
  PM for rooftop units; SLA-backed response windows.
- [Heat pump retrofit](https://example.com/services/retrofit): store-level
  electrification projects with M&V reporting.

## Who we are
- [About Northline](https://example.com/about): leadership, licensing, service area.
- [Safety & compliance](https://example.com/compliance): certifications and COIs.

## Start here
- [Request a site audit](https://example.com/contact): multi-location assessment.

## Optional
- [Maintenance FAQ](https://example.com/faq): contract terms and after-hours coverage.

Notice what is missing: /blog/page/2, marketing adjectives, and duplicate nav labels with no facts attached. The ## Optional heading is the spec’s skip-if-short-context bucket — use it for secondary FAQ and press, not for the commercial door.

How do models and tooling actually use llms.txt?

No major lab’s crawler documentation says “we ingest /llms.txt as a ranking or citation signal.” What vendors document is crawl access and product tokens.

In practice you will see three behaviors:

  1. Direct fetch — some agents, IDE tools, and research workflows request /llms.txt when exploring a domain. The spec was written for that inference-time path.
  2. Indirect use — the same facts appear in HTML that retrieval already prefers. The file keeps your team honest.
  3. No use — many brand-prompt sessions never fetch it. Your HTML truth layer still has to be strong.

Chrome’s Lighthouse audit will flag a server error when it tries to retrieve the file. A 404 is N/A, because providing the file is optional. That is the opposite of “Google Search requires this.” Do not tell a stakeholder that a Lighthouse pass equals ChatGPT citations.

Claim you will hearWhat the primary source actually says
“It’s the official standard”llmstxt.org is a community proposal with a public repo — not RFC-status
“Google requires it”Lighthouse: optional; 404 → N/A
“ChatGPT reads it every time”OpenAI documents GPTBot / OAI-SearchBot / ChatGPT-User, not llms.txt consumption
“It blocks training”Training opt-out is robots.txt / published policy, not this file

So the ROI is dual: better machine briefing when something fetches it, and a canonical outline that improves About/Services copy when you align them. Treat “we shipped llms.txt” as a truth-layer task, not a visibility KPI.

The Answer.AI announcement framed the file for inference — an agent that needs to use a site now — not as a training-corpus contract. Two years later, the v2 page at llmstxt.org still does not claim every chat product loads the file on every brand prompt. Write for the fetchers that exist. Do not staff a program on a guarantee no lab published.

How do I pair llms.txt with robots.txt?

RFC 9309 is the Robots Exclusion Protocol. Google’s robots.txt introduction is explicit: robots.txt manages crawler access; it is not how you hide a page from Google Search. llms.txt does not change that.

Vendor crawlers you actually configure — details in the crawler spoke:

TokenVendor docJobIf you Disallow
GPTBotOpenAI botsFoundation-model training crawlSignals content should not train OpenAI generative models
OAI-SearchBotSameChatGPT search indexSite not shown in ChatGPT search answers (may still appear as nav links)
ChatGPT-UserSameUser-initiated fetchLive fetches may fail; OpenAI notes robots.txt may not apply
ClaudeBotAnthropic crawler articleTraining-oriented collectionSignals future materials should be excluded from training datasets
Claude-SearchBot / Claude-UserSameSearch index / user-directed fetchAnthropic documents reduced search or user-fetch visibility
GooglebotGoogle common crawlersSearch crawlBlocks Search crawl/index path for those URLs
Google-ExtendedSameProduct token for Gemini training / listed groundingDoes not control Search inclusion or ranking

OpenAI states those settings are independent. You can allow OAI-SearchBot and disallow GPTBot. llms.txt cannot express that split. Do not put User-agent: lines in the markdown briefing. Do not assume a beautiful briefing overrides a CDN “block AI crawlers” toggle that 403s the search bot.

Checklist:

  • robots.txt and llms.txt reviewed in the same PR when policy changes
  • Training vs search tokens decided on purpose, not by a single “AI” switch
  • Staging and internal search still disallowed in robots
  • /llms.txt itself is allowed for the agents you want to brief

Google’s common-crawler list is the place to confirm Google-Extended is a product token, not a second Googlebot. OpenAI’s bot page is the place to confirm GPTBot and OAI-SearchBot are independent. Neither page tells you to put those tokens inside llms.txt.

What is the llms.txt implementation checklist?

  1. Inventory the 10 URLs that should settle buyer questions.
  2. Draft the briefing offline; read it aloud — if it sounds like nav labels, rewrite.
  3. Align founding year, HQ, and offer names with About and Organization JSON-LD.
  4. Publish at https://yourdomain.com/llms.txt with a crawlable 200, no auth. text/plain or markdown both work; do not hide it behind a login wall.
  5. Link it from a humans-facing page if you want transparency (footer or AI/info page). Optional: Link: </llms.txt>; rel="describedby" per the spec.
  6. Optional: maintain llms-full.txt or a /docs/llms.txt for longer documentation; keep the root file short.
  7. Re-test a fixed brand-prompt panel after publish; log whether descriptions tighten. One screenshot is anecdote.
  8. Revisit on every pricing/offer change, not on every blog post.

Ship behind a PR that also updates About and Organization schema in the same merge. Split deploys are how llms.txt becomes the only accurate document on the site — which sounds good until HTML retrieval ignores it and quotes the stale About page instead.

What do weak vs strong llms.txt bullets look like?

Weak: - [Services](/services): Everything you need to grow
Strong: - [Outbound sequencing](/services/outbound): Human-approved AI sequences for B2B teams with 5–50 SDRs; HubSpot and Salesforce.

Weak: - [Blog](/blog): Insights
Strong: - [AEO playbook](/blog/answer-engine-optimization-playbook): Full method for AI citations; start here for visibility work.

Weak: - Founded by innovators in 2015ish
Strong: - Founded 2015 in Charleston, SC. Not affiliated with Acme Robotics (Delaware).

The strong versions survive compression. The weak ones become hallucinated mush.

TestFailPass
Read the note without the linkSounds like a nav labelSounds like a fact
Ask “which question?”“Our content”“What do you sell to whom?”
Click the URL301 chain to homepage200 on the claim
Diff vs schema name / foundingDateMismatchSame string

How should llms.txt work for multi-brand and multi-language sites?

If you operate several brands on one domain (unusual but real), separate sections with explicit brand headings and never reuse the same product names across brands without labels. Prefer brand subdomains or distinct domains when the entities are truly separate. The spec’s path-scoped files (/brand-a/llms.txt) help only if the public URLs actually live under that path.

For multilingual sites, either:

  • Publish language-specific briefings only if your stack already localizes that way and you document the convention, or
  • Keep one English canonical briefing that points to localized HTML answer pages

Do not machine-translate the briefing into five languages and leave conflicting founding years in each. Pick a source language for facts.

SetupFile strategyFailure mode
One brand, one domainSingle /llms.txtOver-long Optional section
Two brands, two domainsOne file per hostCopy-paste of the wrong legal name
Two brands, one hostPath-scoped files or labeled H2sMixed entity IDs in one H1
Localized marketing sitesOne fact language + localized HTML linksTwo founding years

Who should own llms.txt on a brand team?

Assign one owner (usually marketing ops or the founder on smaller teams). Sales and PR do not freestyle alternate origin stories. When you launch a new offer, update the HTML page first, then llms.txt, then any PR boilerplate. That order prevents the file from advertising a page that still says the old thing.

TriggerAction
Pricing changeUpdate linked package page first, then briefing bullet
New service lineAdd bullet + deep link; remove if beta and not public
Office moveHQ line + About + schema same day
Executive hire used in salesAdd Person link only if a public bio exists
QuarterlyRead the whole file aloud; cut anything stale

Put the file in the same repo as the site when possible so it ships with deploys. Orphan docs in Notion drift. On material edits, add a “Last reviewed” ISO date at the bottom of the file. Models do not require the date. Your team does — especially when someone asks why a chat still mentions a beta product you removed last month.

Which llms.txt failure modes actually show up?

These are the ways a “we shipped llms.txt” project goes wrong. None of them are theoretical platform lore.

FailureWhat you seeFix
Sitemap pasteModel lists six services you do not sellCut to question-shaped URLs
Briefing ahead of HTMLFile says 2024 rebrand; About still 2019Same-PR update
CDN “block AI”File is 200 for you, 403 for search botsAlign with crawler policy
Auth or noindex on the fileAgents never see itPublic 200, indexable or at least fetchable
Two legal namesChat invents a mergerOne name across file, schema, About
Empty lawyer fileStakeholders think the work is doneEither approve facts or do not publish
Treating Lighthouse as AEOPass/fail used as a citation KPILighthouse optional ≠ engine honor

The expensive one is the split deploy: sales starts quoting the new package from the briefing while Google and chat still retrieve last year’s HTML. Bravery is not a restore strategy. Same merge, or do not ship the file.

How do I QA llms.txt before I call it done?

  1. Fetch production with curl -I https://yourdomain.com/llms.txt — 200, not 401/403, not a homepage HTML fallback.
  2. Confirm Content-Type is text (text/plain or text/markdown), not an HTML error wrapper.
  3. Paste the file into a blank chat and ask: “Summarize this company in three sentences.” If the summary invents scope, your briefing is vague.
  4. Click every link. Dead links teach machines to ignore you.
  5. Diff against Organization schema name, url, foundingDate, areaServed.
  6. Run a fixed 10-prompt brand panel and note whether descriptions tighten over two weeks — not whether one session “felt better.”
  7. Re-run Chrome Lighthouse’s llms.txt check if you care about the agentic-browsing category. A 5xx is a fail; a missing file is still optional.
  • H1 is the real name
  • Blockquote is facts, not a tagline
  • Every H2 list item has a live URL
  • Optional section is actually optional
  • robots.txt policy matches the story you tell in the briefing

How do I draft llms.txt in a 45-minute workshop?

Block a calendar slot with the founder (or whoever can bind facts) and a marketer who knows the URL map. Whiteboard only four columns: Claim, Evidence URL, Owner, Public?. Fill ten rows max. Anything without an evidence URL dies. Anything not public dies. Rewrite the survivors into markdown bullets with links.

Then do a hostile edit pass: delete every adjective that does not change a decision. “Trusted,” “innovative,” and “full-service” almost never survive. Numbers, certifications, ICP boundaries, and geography do.

Finally, paste the draft into a blank model chat and ask three prompts: (1) What does this company sell? (2) Who should not hire them? (3) Which URL should I read first? If the model invents a fourth product line, your briefing leaked ambiguity. Fix the markdown before it ever hits production.

What are common stakeholder objections to llms.txt?

“Nobody reads plain text files.” Humans might not. Some agents and docs tooling do — and the drafting exercise improves human pages regardless.

“We will wait for an official standard.” Waiting is a decision to stay ambiguous. The briefing pattern is useful even if filenames evolve. Do not wait for IETF status that this proposal does not have.

“Our lawyers want the file empty.” Empty is fine if counsel blocks public claims; then invest in HTML pages they will approve. Do not publish a teasing file full of hedges that teach nothing.

“We already have a sitemap.” Sitemaps list URLs. Briefings state truths. Different jobs. See sitemaps.org.

“Lighthouse failed.” Read the audit: server errors fail; absence is N/A. Fix 5xx. Do not treat N/A as an emergency.

“This will get us cited.” Citations come from retrievable, consistent, corroborated pages — the stack in the AEO playbook. The file is one layer.

FAQ

What is llms.txt?

It is a voluntary markdown file, usually at /llms.txt, that summarizes the organization, offers, and key URLs for language-model systems and related tooling. Jeremy Howard proposed it in September 2024; the living format is at llmstxt.org. Think briefing document, not sitemap clone, and not a formal internet standard.

How do I write llms.txt for a business?

State who you are, who you serve, what you offer, and link to the pages that prove each point. Use the spec order: H1 name, blockquote, then H2 lists of [name](url): notes. Keep it short, factual, and consistent with your About page and schema. Skip blog dumps and hype adjectives.

Does every AI model read llms.txt?

No. Support varies by product, tool, and session. Chrome Lighthouse treats a missing file as optional. Vendor crawler docs describe robots.txt tokens, not a guarantee that brand prompts load /llms.txt. Publish it as part of a broader AEO truth layer, and make sure the same facts exist in HTML.

Should llms.txt block AI training?

That is not its job. Use robots.txt, meta rules, and your published terms for crawler and training preferences — OpenAI’s bot overview and Anthropic’s crawler article are the starting points. Use llms.txt to clarify facts for systems that request a summary.

How long should the file be?

Long enough to disambiguate the brand and point to answer pages — often 40–120 lines for a focused company. If you need a manual, use llms-full.txt, a path-scoped /docs/llms.txt, or documentation URLs. If you need a URL inventory, that is the sitemap.

Can llms.txt fix hallucinated brand facts alone?

Rarely alone. It helps when a tool fetches it, but you still need consistent HTML, schema, and off-site corroboration. A briefing that contradicts About is worse than no file. Pair it with the rest of the AEO playbook.

CTA

Ship a briefing, not a URL landfill. When llms.txt matches your schema and canonical pages, you give answer engines fewer reasons to invent you.

For the full system — entities, content clusters, citation measurement — read the AEO playbook. To have Spurlock Studios baseline your file and the rest of the truth layer, start at /visibility or book a visibility audit.

FAQ

What questions does this article answer?

What is llms.txt?
It is a voluntary markdown file, usually at `/llms.txt`, that summarizes the organization, offers, and key URLs for language-model systems and related tooling. Jeremy Howard proposed it in September 2024; the living format is at [llmstxt.org](https://llmstxt.org/). Think briefing document, not sitemap clone, and not a formal internet standard.
How do I write llms.txt for a business?
State who you are, who you serve, what you offer, and link to the pages that prove each point. Use the spec order: H1 name, blockquote, then H2 lists of `[name](url):` notes. Keep it short, factual, and consistent with your About page and schema. Skip blog dumps and hype adjectives.
Does every AI model read llms.txt?
No. Support varies by product, tool, and session. Chrome Lighthouse treats a missing file as optional. Vendor crawler docs describe robots.txt tokens, not a guarantee that brand prompts load `/llms.txt`. Publish it as part of a broader AEO truth layer, and make sure the same facts exist in HTML.
Should llms.txt block AI training?
That is not its job. Use robots.txt, meta rules, and your published terms for crawler and training preferences — OpenAI’s [bot overview](https://developers.openai.com/api/docs/bots) and Anthropic’s [crawler article](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) are the starting points. Use `llms.txt` to clarify facts for systems that request a summary.
How long should the file be?
Long enough to disambiguate the brand and point to answer pages — often 40–120 lines for a focused company. If you need a manual, use `llms-full.txt`, a path-scoped `/docs/llms.txt`, or documentation URLs. If you need a URL inventory, that is the sitemap.
Can llms.txt fix hallucinated brand facts alone?
Rarely alone. It helps when a tool fetches it, but you still need consistent HTML, schema, and off-site corroboration. A briefing that contradicts About is worse than no file. Pair it with the rest of the [AEO playbook](/blog/answer-engine-optimization-playbook).
Sources

Last reviewed — llmstxt.org v2 spec, Answer.AI proposal, OpenAI / Anthropic / Google crawler docs, RFC 9309, sitemap protocol, and Chrome Lighthouse llms.txt audit checked 2026-08-16.

More from this lane

AI Visibility

All →
Book the audit