Skip to content

Build1 publisher3 min readPublished

AI-written code fails the same four ways, and every gate you own reports green

A scan of AI-built repositories reports the same patterns everywhere: swallowed errors, defaults standing in for real data. All of it compiled, linted clean and passed the tests that existed.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying AI-written code fails the same four ways, and every gate you own reports green
Generated illustration

What happened

  • The authors examined vibe-coded GitHub repositories identified by a "fingerprint" (projects built primarily with AI coding assistance) across a range of frameworks, languages and team sizes, to find patterns, failure modes and recurring structural problems.
  • The authors expected variation across different tools, teams and codebases, but report getting the same handful of patterns over and over regardless of any of that.
  • The code compiled, the tests passed (when there were tests at all), and the linters were clean, while underneath errors were being swallowed silently, missing data was papered over with defaults, and the codebase was quietly rotting from the inside.
  • Silent failures were the most dangerous and most common pattern found: an AI assistant writes a try/catch, the catch block logs the error or does not, and execution continues as if nothing happened, so the app looks like it is working while data is wrong, an operation got skipped, and the system is in a state nobody accounted for.
  • The proposed fix is that every error must be visible in three places at once: the console, the UI, and as a thrown exception that actually halts execution; an error that only shows up in a log nobody is watching is not handled, it is hidden.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A team writing on dev.to says it went looking for structural patterns in GitHub repositories it had identified as AI-built, using a "fingerprint", across a range of frameworks, languages and team sizes [1]. It expected the problems to vary by tool, team and codebase, and reports finding the same handful of failure modes regardless [2]. The consequence is not really about AI. According to the post, the code compiled, the linters were clean and the tests passed where tests existed at all, while errors were being swallowed silently and missing data was papered over with defaults [3]. Every automated gate in a normal pipeline reported green on code the authors describe as quietly rotting. Two patterns carry most of that weight. The first is the silent failure: an assistant writes a try/catch, the catch logs the error or does not, and execution continues as though nothing happened, so the app looks fine while an operation has been skipped and the system is in a state nobody accounted for [4]. The second is what the post calls phantom correctness, code that compiles and lints and then operates on the wrong data or a fabricated value [6]. Its two biggest contributors are named specifically: `any` types used to make a type error go away rather than resolve it, and hardcoded values standing in for something that should have come from a real data source [7]. Both look completely fine in a diff, and neither throws [8]. That is the mechanism worth internalising. A compiler checks that types are consistent, and `any` is a consistent type. A linter checks form. A test checks the assertion someone wrote. None of them can tell the difference between a value that came from a data source and a value that was typed in to make the screen render. The post's framing is that this is not an intelligence problem: models are optimised to produce code that looks correct, and nothing in the default generation loop checks the separate claim that it is correct [9]. It also notes a behavioural version of the same gap, where an assistant declares a task done once the obvious high-priority wins are resolved and defers warnings and info-level findings, which in practice means never, because nothing forces a return trip [10]. The suggested remedy is that every error should be visible in three places at once: the console, the UI, and as a thrown exception that halts execution [5]. Treat the third part as a debating point rather than a rule. Halting is one way to make failure impossible to ignore, and in plenty of production paths it is the wrong one; the defensible core of the advice is that an error which only lands in a log nobody watches is hidden, not handled [5]. The post is candid that authoring-time guardrails such as rules and skills work, but only at the moment code is written [11]. It lists a second class that appears after execution: the happy-path trap in async code, meaning the second concurrent click rather than the first; vendor SDKs imported directly across dozens of files; tests that assert implementation details and break on every refactor [12]. These surface in production, or six months later when someone who was not there has to touch the code [13]. And because generation is probabilistic, it argues code drifts from design intent regardless: a prop renamed in a technically valid way that breaks a contract, a spacing value approximated instead of read from a token, a component restructured so it works but no longer matches the design system [14]. Read the caveats. The findings end in a governance framework of six principles plus forbidden and required controls, organised by class of problem rather than technology, designed to drop into whatever you already use [15], which makes this a diagnosis with a product attached.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories