Skip to content

Build1 publisher3 min readPublished

Most llms.txt files on GitHub put prose where the format expects link lists

Just 30.6% of 438 llms.txt files on GitHub follow the proposal's full structure, though 96.6% open with its required H1, a dev.to count found. A check that stops at the header passes files whose body sections a link-reading program cannot rely on.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • In 231 files, or 52.7%, at least one H2 section holds paragraphs of prose where the proposal describes a list of links.
  • 84 files, or 19.2%, use H3 or deeper headings or a second H1, structures the format does not describe.
  • Of 56,557 full web addresses in the files, 6,170, or 10.9%, point straight at a .md, .mdx or .txt file.
  • 119 files, or 27.2%, link mostly by relative path, and those links resolve against wherever the file is fetched from.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Tools that pull links by walking H2 lists alone cannot treat a published llms.txt as dependable input; they need a fallback that reads prose sections or treats the file as plain markdown.
  • decision A validator that checks only the H1 and blockquote approves nearly every file, so teams shipping llms.txt need a per-section test of the H2 body before calling the file conformant.
  • exposure Files built on relative links work only when served from the site they describe; a copy fetched from the repository or anywhere else points its links at the wrong base.

The proposal splits everything after the summary into two parts [10]. First come sections of any kind except headings. Then come H2 sections, each a "file list": a markdown list whose items are a [name](url) link, with optional notes after a colon [10]. Files from 331 repositories were checked against v2 of that text as it read on October 4, 2026 [1].

The list-only figure needs one correction. Nineteen of the files that keep every H2 section a link list pass because they have no H2 sections at all [5]. Strip those out and 124 files keep every section a list, out of the 419 that use H2 headings [8]. Among files that use the section structure, 29.6% use it as specified [20].

Agents still get something. According to the count, the file is markdown either way, and the proposal expects agents to view or search it and follow its links [11]. What drifts, the count's author wrote, is "the predictability the format promises, that a program can find the links by reading the H2 lists alone" [11]. A program built on that promise skips any section written as prose. Some files give it nothing to follow: 102 (23.3%) hold no links at all, while the median file holds 13 [9].

The proposal also asks that links point at content a language model can read easily [12]. It suggests a markdown copy of each page at the same address, with .md appended or swapped in for the extension [12]. The author calls the address count an undercount. Version 2 lets a page point to its markdown version through a link relation, and counting addresses cannot see that [14].

The sample is repositories. The proposal is written for files served on websites, and the count read files as committed to GitHub [18]. Only 159 sit at a repository root; the others are in subdirectories, for example the public folder of a documentation site, from which a build step puts them online [7]. For the percentages to describe what crawlers fetch, those builds would have to ship the files unchanged. The count did not read served copies [18].

The author's check for the body is one line of awk [16]:

``` awk '/^#/{s=($0 ~ /^## /)?$0:""; next} s && NF && !/^[-*+] \[/ && !/^[ \t]/{print NR": "s": "$0; s=""}' llms.txt ```

On a heading line it keeps the heading if it is an H2 and clears it if not. Inside an H2 section, the first non-blank line that is neither a list item opening with a link nor an indented continuation is printed with its line number and heading [16]. Run on a scratch file, it printed one line: "18: ## Pricing: Acme charges per request, billed monthly." [16]. The pattern requires the link to come first in the item, so "- See [docs](url)" is flagged [21]. That is strict, and it matches the item shape the proposal describes [10].

It reads every line as text, so a ## line inside a fenced code block counts as a heading [17]. The author's own count ignored headings inside fences and dropped a leading byte-order mark [19]. I think this is good work for one line: it tests the rule files actually break, and its author documents where it misreads.

What to watch

  • A count of llms.txt files as served from websites instead of as committed to GitHub, to test whether builds ship what the repositories hold.
  • A revision of the llms.txt proposal beyond v2 that either describes prose sections and H3 headings or tightens the file-list rule.
  • Whether the share of links reaching markdown rises once link-relation pointers are counted alongside .md addresses.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories