Skip to content

Build1 publisher3 min readPublished

Writing an llms.txt forces a business to state what it refuses to claim

Jeremy Howard's convention has been public since September 2024, and no large model provider has confirmed reading the file, so the case for spending an hour on one rests on what the drafting forces you to decide.

The Engineer · Build desk

Illustration accompanying Writing an llms.txt forces a business to state what it refuses to claim

What happened

  • The llms.txt convention, a plain text summary of a site written for language models rather than crawlers, was published by Jeremy Howard in September 2024 as a proposal and not a standard.
  • By the dev.to post's account, a few thousand sites have adopted it and nobody who runs a large model has said whether they use it.
  • The convention asks for Markdown at /llms.txt, an H1 naming the business and a blockquote summary, and leaves everything after that unspecified.
  • The post sets out six elements that earn a place in the file, among them the boundary of what the business does, alternative spellings of names, and explicit citation terms.
  • It names the failure mode as writing the file as a pitch: adjectives true of every competitor make a model summarising ten businesses produce ten identical sentences and choose on some other basis.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Citation terms are a choice most sites currently make by default and by omission; writing the file puts a length and an attribution requirement on the record, or documents that you never decided.
  • constraint Anything that fetches raw HTML cannot use a contact form assembled by JavaScript, so a plain address in the file is the only route such a system can repeat back.
  • cost The spend is an hour of drafting plus upkeep on a prose duplicate, set against an unconfirmed retrieval benefit. The internal clarity is the one part of the return you can verify.
  • capability A written boundary lets a model exclude you from queries you do not serve. Structured markup cannot do that, because refusals have no fixed shape.

The format specifies almost nothing. Markdown served at /llms.txt, an H1 that names the business, a blockquote that summarises it, and after that, by the post's own account, the convention goes loose [5].

So everything useful in the file comes from what you decide to put in it. The post sets a budget: fewer than a thousand words to state what the business does, who it is for, who it is not for, what it charges, and what it refuses to claim [6]. It then names six elements to include [7]. Split a thousand words six ways and each element gets about 167 [21]. One short paragraph per decision.

The expensive one to write is the boundary. The post's worked example is a single sentence: "this is not a general agency and does not sell rankings, packages or volume". That tells a model more than six bullet points of capability, the post argues, because it can be used to rule you out [9]. The author wrote: "Being ruled out correctly is not a loss. It is the only way being ruled in means anything." [10]

Two of the six are ordinary data hygiene. Write the unaccented spelling of a founder's name alongside the accented one and say they refer to the same person. A system then resolves you to one entity instead of two partial ones that never accumulate [11]. Publish a contact address a machine can read and repeat, because on most platforms the contact form is drawn by JavaScript, and a machine reading the page may find no route to you at all [13]. That claim is checkable in a minute: fetch your own page without executing scripts and see whether an address survives.

The post also draws a line between this file and structured markup. Schema handles facts with a fixed shape and does nothing for judgement, scope or refusals. Of the four files that describe a site to machines, only one is prose [14][15].

On numbers, the instruction is narrow and I would keep it. A percentage with no method attached is worse than no percentage. If a number matters, say how it was measured; if it cannot be measured honestly with the tools you have, say that instead [17].

The post presents no evidence that publishing the file changes what any assistant says about a business, and it says as much: no major AI company has confirmed that it reads one [3]. What it argues instead is that the drafting is the return, and that businesses which have written this down tend to have a clearer website because the same clarity reaches every page [19]. It poses the question itself: is the hour worth spending, given that assistants might not read it? The answer it gives is yes [20]. The summary line is the author's: "The file is the cheap part. The decision is the asset." [18]

In my view that holds for the six decisions, which you can check against your own pricing page today. It does not hold for the path. A hand-written prose file is a second copy of facts that live elsewhere, and that copy will drift from the pages it came from.

What to watch

  • Any crawler or assistant documentation from a large model provider that names /llms.txt as a path it fetches.
  • A measured count of sites serving /llms.txt, and any test showing whether the file changes what an assistant answers.
  • Whether the loose half of the convention acquires a schema or a validator. That would turn drafting judgement into a lint check.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories