Skip to content

Build1 publisher2 min readPublished

Code-writing LLMs repeat invented package names often enough for squatters to register them first

USENIX Security 2025 researchers found 19.7% of packages suggested by 16 LLMs were fake, and 43% of those names recurred on every re-run. Names that repeat can be registered ahead of time, so a team has to vet a suggested dependency before installing it, even when the install succeeds.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Code-writing LLMs repeat invented package names often enough for squatters to register them first
Photo: lasso.security

What happened

  • The USENIX paper, titled "We Have a Package for You!", generated 576,000 code samples and checked every package the models recommended.
  • Open-source models invented packages at rates up to about 22%, while GPT-4 Turbo was the best commercial performer at 3.59%.
  • A 2026 re-evaluation of newer frontier models found hallucination rates between roughly 4.6% and 6.1%.
  • A separate 2026 cross-model study found 127 package names that five different frontier models all invented identically.
  • Measured by edit distance, only about 13% of the invented names were simple typos of real packages, and nearly half resembled nothing that exists.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint String-similarity scanners are tuned for typos, and about 87% of invented names are not simple typos, so most of these packages get past the filters teams already run.
  • exposure The target is the developer who copies a suggested install line exactly; typing care, the usual defence against typosquatting, does not help here.
  • contradiction The post calls the 2026 figures better, yet their 4.6% floor sits about a point above GPT-4 Turbo's 3.59% in the 2025 study, so the gain came at the bad end of the range.

An invented package name fails safe on its own. A dev.to post that summarises the research says so directly: without an attacker, the model's mistake ends in a failed install [12]. The harm comes from a second party who registers the invented name first and puts malicious code inside [11]. Seth Larson of the Python Software Foundation named the attack slopsquatting in 2025 [1].

Registering a fake name only pays if the model asks for it again. According to the post, the USENIX researchers logged 205,474 unique names that did not exist [4]. Nobody registers that many in the hope of a hit, and the post argues that random hallucinations would protect developers for that reason [15]. The re-run test takes that protection away. The researchers took 500 prompts that had produced a fake package and ran each one ten more times [7]. Names that came back on every run give an attacker a list to register. The post lays out the routine: run popular models against popular prompts, collect the names that recur, register them with malware, and wait [11]. "The model does the targeting for them, for free, every time a new developer asks a similar question," the post says [16].

Existing typosquat detection was built for a different mistake. It measures string similarity by counting the single-character edits between a new name and a known one [10]. Names like aws-helper-sdk and fastapi-middleware sound plausible without being close to any real package, so the post argues they get past that check [13].

The headline rate describes someone else's workload. It pools every model the study tested, open-source ones included [2][5]. A team on a current frontier model should start from the 2026 range instead [6]. That comes to between four and six invented names in every hundred suggestions [3]. Even that figure transfers only if the team's prompts look like the ones the researchers used. The recurrence figure covers less than it seems to. It was measured only on prompts already known to produce a fake name, so it tells you how stable a hallucination is once it has appeared [7].

In my view, an install line from a model is untrusted input. Asking the registry whether the name exists does not help, because once the attacker has uploaded the package the answer is yes [11]. The check I would trust compares the name against a record the model did not write: the dependencies a person on the team already chose. Any name outside that set waits until someone has read its package page and its code.

What to watch

  • Confirmed malicious PyPI uploads under names that frontier models are documented to invent repeatedly.
  • Publication of the prompts behind the 2026 re-evaluation, so teams can test whether the 4.6% to 6.1% range holds for their own work.
  • Typosquat scanners adding checks that do not depend on edit distance to a known package name.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories