Build1 distinct publisher2 min readPublished
The file exists so an agent can skip your documentation and read the install steps instead. Pandex checked 8,565 of them, found 237 references pointing at names anyone could register, and registered one.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Here is what actually happens. An agent reaches a docs site and fetches llms.txt, because parsing the whole documentation tree burns tokens and context window; the file hands over the language, the environment, the dependencies and often precise setup and installation commands in a few hundred words [1]. The agent then treats those commands as the way to run the product. The illustration in the report is the entire bug: the docs say `pip install wtf-software` because that is what the package was going to be called, the shipped package is `wtf-software-beans`, and someone else owns the shorter name [11]. The convention descends from robots.txt, which guided crawlers to content and is still widely used [2]. The difference that matters is that this file is read by something holding a shell.
237 references across 8,565 files is one dead reference for every 36 files [12]. Note that the count is references, not files, so the share of files carrying at least one is at most 2.8 percent [12]. The names were spread across six package registries plus expired .dev and .io domains and abandoned Render, Vercel, Fly and Netlify subdomains [5], which means no single registry's namespace policy closes this.
The model numbers are the part I would not carry over unexamined. Pandex reports GPT-5 Luna and Sol running the payload 90 percent of the time or more, against Claude Opus 4.8 at medium effort at 30 percent [9], a threefold spread [13] that the researchers attribute to how autonomous the frontier models are [9]. For that ratio to describe your deployment, your agent would need to install and run packages without an approval step, in an environment where callback traffic actually leaves, prompted as loosely as their one line was [7][8]. Change any of those and the number moves. The direction survives the caveats: the more of the decision you hand to the agent, the more the docs it reads become code it runs.
The detail that settles it for me is that Pandex found a case where someone had already pulled this off with real malware, and notified the publisher [10]. The technique was in production before the write-up. The calendar-invitation attack on Gemini a year ago made the same point through a different container [14].
So if you publish llms.txt, you are publishing a dependency manifest. It has no lockfile, and no CI job fails when one of its names stops resolving to you. The work that follows is dull: every package name and URL in that file, reconciled against what you actually ship, on whatever cadence you already use for dependency review.
Ranked by verification strength, evidence, and original report placement.
llms.txt is often hosted on a software product's website and contains a brief description, setup instructions and quick installation steps, so that a bot can read it instead of spending tokens and context window parsing the whole documentation; it tells the agent what language the code uses, the environment it runs in, any dependencies, and often precise setup/installation instructions.
Sites began publishing robots.txt to guide search bots to content, it is still widely used today, and it has now been supplemented by llms.txt, a file containing textual instructions for AI agents to follow.
Researchers from Pandex got their own code to run on AI agents belonging to companies in the Fortune 500, via an llms.txt guidance file.
Across 8,565 llms.txt files checked, the researchers found 237 references to software packages that no longer exist, do not exist yet, are mistyped, are now hosted elsewhere, or imply out-of-date information compared with the current documentation.
According to Pandex, the packages spanned PyPI, npm, RubyGems, NuGet, crates.io and Packagist, and domains ranged from expired .dev and .io registrations to abandoned Render, Vercel, Fly and Netlify subdomains, all free to the first person who clicks 'claim'.
Pandex created their own Python and Node 'malware' that would call back home and sit waiting; four minutes after going live, there was a bite on the hook.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
science
A backdoored litellm release turns every CI job that installed it into a credential incident1 distinct publisher
build
Every one of thirteen named 2025-26 incidents ran on a credential that still worked1 distinct publisher
build
Three files, three contracts: robots.txt, sitemap.xml and llms.txt are not rivals1 distinct publisher
science
TeamPCP hid its infostealer inside the scanners that audit everyone else's code1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one firm's unpublished-here experiment
Everything load-carrying in this story — 8,565 files, 237 claimable references, four minutes to first execution, 90 percent versus 30 percent — reaches us through Tom's Hardware quoting Pandex, and Pandex's own write-up is not in front of us. No victim, package, domain or registry entry is named, no trial counts accompany the model rates, and nothing has been checked by a second party. The mechanism itself is old and cheap to believe; the specifics are a single retelling.
The files are everywhere; the exposure is unmeasured
That 8,565 llms.txt files were simply sitting there to be scanned is the most solid uptake fact in the story — agent-facing instruction files are now ordinary furniture on product sites. The other half of the equation is blank. Nobody counts how many agents read these files, in whose environments, with what permissions. One anonymous agent biting inside four minutes shows the path is live; it says nothing about how wide it is.
'Easily trick' outruns 237 in 8,565
The framing does more work than the data under it: 237 claimable references across 8,565 files is one file in thirty-six, and the Fortune 500 demonstration involves targets nobody will name. Yet the underplayed part is the direction of travel — the more autonomous the model, the more often it ran a stranger's setup steps, and the researchers needed no injection, no links and no social engineering to get there. Today's blast radius is oversold; tomorrow's is undersold.
The firm that found the hole also sizes it
Pandex discovered the gap, registered the name, ran the experiment and gets to characterise how frightening the result is — security research that lands on the Fortune 500 doubles as a shop window, and we are given no view of what the firm sells. Tom's Hardware supplies its own gradient with 'easily trick' and a data-is-code sermon rather than pressing on disclosure or provenance. Neither incentive is concealed or unusual; the point is that no one in this story has any reason to talk the finding down.
Shape solid, magnitude unsettled
Two different questions get two different answers here. Can an agent be led to execute code from a claimed name in a documentation file? Believably yes — the pathway is mundane and consistent with the calendar-invite and MCP results the piece recalls. How much of the installed base is exposed right now, and to whom? On that, one outlet relaying one firm's first pass, with no named parties, is thin. Act on the mechanism; hold the numbers loosely until someone else counts.