Build1 distinct publisher3 min readPublished
Dan Luu's August 30th essay argues engineers automate their own workarounds until the underlying bug stops producing reports, which matters more now that agents can aim a lot of cheap code at the wrong problem.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
MCP's roadmap fast-tracks five priorities and quietly queues everything else1 distinct publisher
invest
Microsoft's Idle AI Chips Are A Construction Problem, Not A Shortage1 distinct publisher
build
A 7% ripgrep win is the wrong number in Dan Luu's agent experiment1 distinct publisher
build
Perf work stopped being a specialist queue item, and slow endpoints became a choice1 distinct publisher
A bug tracker runs on surprise. Someone hits behaviour they did not expect, notices that they noticed, and writes it down. Habituation removes the first step. The pause before renaming a fresh Google Doc, learned because text typed immediately could be overwritten, stops being a decision and becomes part of how you type [3]. Luu's Microsoft example is the same mechanism with worse plumbing: staff disabled Wi-Fi before logging in, because an unreliable service rejected the login with a server-unavailable error and a disconnected laptop skipped the check entirely [4]. A login check that passes when the network is gone is at least consistent about what it is checking.
The defect keeps firing at the same rate for everyone else. What changes is the reporting path. And because the internal quality signal is a set of humans noticing things, a team whose novelty detector has adapted to its own product reads silence as quality [2]. Luu says he has watched teams ship products with little chance of success because everyone involved believed the quality was high [9].
He also offers a cheap instrument. When executives asked him to evaluate products, the troubling cases were the ones requiring several non-obvious workarounds while internal discussion described the thing as working well [7]. That gap is countable. Nobody needs a research budget to count the steps in their own happy path that a new hire had to be told about.
The numbers get uncomfortable at scale. Luu puts his own rate at hundreds of worked-around bugs in an average week and thousands in a bad one, an anecdotal figure he previously backed with a published log of one week of newly encountered bugs [5]. Take the low end of "hundreds" as 200 and a 40-hour week, and that is five per hour, one roughly every twelve minutes [15]. Most of those never became a sentence anyone else could read.
The credential matters here only because it locates the claim. Luu spent eight years at Centaur Technology on x86 processors and validation, was the second engineer on the project that became Google's TPU, and later worked on Microsoft's BitFunnel [6]. This is a debugging observation, not a theory imported from usability research.
Which is why dogfounding inside a startup measures the wrong population. Engineers are unusually good at routing around software problems and then restating the route as reasonable instructions [8]. The output of that test is a statement about the tester's skill. Deliberately trading polish for speed is a legitimate product decision, but it only stays defensible for a team that can still see the trade it made [9].
Agents change the cost structure on one side only. They generate features, tests and patches quickly, and the team still has to say which behaviour is broken and whether a proposed fix improved the experience [10]. Luu's proposed remedy, LLMs standing in for different types of ordinary user to check whether an issue reproduces across scenarios, is the most testable part of the essay [11]. For that number to transfer to your product, the model has to be denied exactly the context your team has internalised: no primed instructions, no internal build notes, no ordering hints. Prompt it with the team's own workaround and you have rebuilt the insider in software. Luu concedes model-driven testing cannot supply the missing judgment by itself [12], and that judgment is exactly what no agent can supply on its own.
Ranked by verification strength, evidence, and original report placement.
Dan Luu published an essay titled "Bug Blindness" on August 30th diagnosing why software teams ship products users struggle to operate: the people building them have learned to stop seeing their own bugs.
After opening a new Google Doc, Luu learned to wait before changing its title because text entered immediately could be overwritten; he also developed timing habits around search inside Google Docs, where invoking search at the wrong moment could open the browser's native search instead.
At Microsoft, Luu says he and other employees sometimes disabled Wi-Fi before logging into their laptops: an unreliable service could reject a login with an error about unavailable servers, while disconnecting the laptop caused the check to be bypassed.
Luu worked eight years at Centaur Technology on x86 processors and related validation, then became the second engineer on the project that developed into Google's Tensor Processing Unit, and later worked on Microsoft's BitFunnel search infrastructure; BitFunnel's team page describes earlier work across network virtualization hardware, deep-learning hardware and x86 and ARM processors.
Luu says executives at several employers asked him to evaluate products when they wanted an opinion from someone likely to identify problems and push for fixes; the troubling cases were products that required several non-obvious workarounds while internal discussion described them as working well.
Luu traces his awareness of the pattern to childhood: a mechanical mouse had accumulated enough dirt to move erratically, and he had unconsciously learned to compensate by throwing his hand in countervailing directions, while a friend found the machine nearly impossible to use.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 29, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One engineer's memory, relayed once
Almost everything rests on Luu describing his own habits, and Runtimewire relaying him without a second voice. The strongest items are checkable in principle by anyone with the products — the Google Docs title overwrite, the login that succeeds only offline — yet nobody in this reporting has tried them, and Google and Microsoft are not asked. The one detail with an outside trace is his résumé, via the BitFunnel team page. Internally consistent and vividly specific; entirely unaudited.
Nothing to count yet
This is an argument, not a thing in use. No team says it changed a release check because of it, no product ships the fresh-user observation Luu recommends, and the idea of running language models as stand-in users appears strictly in the conditional — could give small teams cheap unfamiliar behaviour, not did. Until someone reports trying it, there is no uptake to score.
Careful, except at the hinge
Runtimewire is disciplined where it would be easiest not to be: it calls the bug tally anecdotal in the same breath as reporting it, and it gives a whole section to what a test generator cannot do — force a team to admit a familiar click sequence is unreasonable. The overreach sits at the joint. Moving from a mechanical mouse and a doc-title race to 'AI raises the value of teams that recognise broken workflows before customers leave' is a market thesis carried on anecdote, with no cost curve and no churn figure to weld the halves together.
No one is selling anything
The pressures here are mild and mostly point away from flattery: Luu names defects in products at Google and Microsoft, both former employers, and the story mentions no tool, subscription or funding attached to the argument. What remains is the pull of timeliness. Luu had sat on this thesis for about a decade, and it is the coding-agent frame that makes it publishable now — which is exactly the frame Runtimewire leads with, in its headline and its opening line about why it matters.
Sure what was said, unsure it generalises
We can vouch that the essay says what Runtimewire says it says; the reporting is close to its source and flags its own weak points. Beyond that we are guessing with the author. One publisher means nothing to triangulate against, the central pattern rests on a witness reporting on his own perception, and the forward-looking half — agents multiplying misdirected fixes, models standing in for confused users — has no observed instance behind it.