Build1 distinct publisher3 min readPublished
A dev.to writeup argues the failure is missing memory across sessions, not code quality per turn. The remedy it ships is 63 repository artifacts, and no measured before-and-after.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Each of the three defects named in the post turns on timing rather than syntax: a rate limiter whose check and consume were not atomic at the call site, a heartbeat that proved a worker was pinging rather than progressing, and a held database transaction that outlived the request that opened it [5]. An agent that writes the code and then writes the test for that code will pass all three [6], because the test inherits the same assumption about ordering that the code does. The observation worth keeping is not that agents write bad lines. It is that the repair does not travel: the patch lands in the file where someone noticed the symptom, and the class of defect has nowhere to live unless something outside the conversation records it as a standing rule [7].
That is why a stronger model is the wrong purchase for this particular problem. The author's own position is that a frontier model in 2026 writes fine syntax all day [13], and the bugs he documents were syntactically clean and green on their own tests [6]. Diff review, which is the thing most teams have actually instrumented, scores each turn in isolation and would have signed off on every one of them.
The remedy carries a price the post states in passing. LEO ships 41 codified laws and 22 specialist roles [1], which is 63 artifacts per repository that somebody has to write, keep current, and keep true [1]. The count is less interesting than the bookkeeping around it: the project's changelog logs every rule the system has ever added together with the reason it was added [4]. My read is that the reason field is the part that survives contact with a team, because a rule nobody can justify gets dropped the first time it is inconvenient, and a rule file that has quietly gone wrong instructs the agent with exactly the same authority as one that is still correct.
Be clear about the evidence. It is one practitioner's account: paid engagements on multi-tenant SaaS platforms, one of them with background AI pipelines [3], plus his own project's rule changelog [4]. No before-and-after defect counts appear, and no second team runs the rules against a codebase the author did not write [2]. The claim that the laws and roles stopped the recurrence is his [1], and it is the kind of claim only a comparison can carry.
The argument that generalises, whether or not LEO itself is any good, is about location. The author insists the protocols sit in files inside the repository the agent is working in, not in a paragraph of best practices somewhere in training data and not on a vendor's central server the team cannot see into [12]. Rules in the repo get reviewed, blamed and reverted on the same machinery as code. Memory held on someone else's server is a dependency you cannot diff.
Ranked by verification strength, evidence, and original report placement.
A developer writing on dev.to says his autonomous coding agent's problem is missing memory rather than code quality, and is open-sourcing the rule set as LEO: 41 codified laws, 22 specialist roles, and a file-based memory system that he says stopped the agent from quietly re-breaking the same production defect every few weeks.
The author calls this more expensive than any single wrong line of code because the defect recurs on a schedule instead of getting fixed once.
The author says he directed AI coding agents on real paying engagements, including multi-tenant SaaS platforms, one of them with background AI pipelines.
The project's changelog file roles/SYSTEM_UPGRADE_MANIFEST.md logs every rule the system has ever added, with the reason for it.
The documented recurring defects were: a rate limiter that could be starved by its own retries because check-and-consume was not atomic at the point of call; a background worker whose heartbeat proved it was pinging rather than making progress; and a held database transaction that outlived the request that opened it and held locks until something else timed out.
The author states that in each case the agent's code was syntactically perfect and passed its own tests.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-published practitioner account
One source, self-authored, with the author's own project changelog as the only documentary record. The failure modes are described concretely enough to be recognisable, which is why this is not near zero, but there are no defect counts, no before-and-after, no control, no second team, and no independent corroboration of the claim that LEO stopped the recurrence.
Announcement plus author's own use
The only observable uptake is the author announcing the open-sourcing and disclosing that he used agents under these rules on his own paying engagements. No external adopters, repository metrics, license, downloads, or third-party deployments appear in the supplied source.
Sweeping framing on anecdotal support
The rhetoric runs well ahead of the support: architecture-level amnesia, a defect 'far more expensive than any single wrong line of code', and a 63-artifact remedy credited with ending recurrence — all resting on one author's unquantified experience. The gap is not larger because the underlying failure taxonomy is specific and plausible rather than vague, and the post does concede that per-turn code quality is not the issue rather than overselling model deficiency.
Author promoting his own release
The piece is a launch narrative for the author's own open-source project, published on a developer-community platform where the same post serves as distribution. The diagnosis and the product are authored by the same person, and the post's success criterion — adoption of LEO — is the thing its evidence is meant to establish. This is disclosed openly in the standfirst, which is why it is not scored higher.
Clear on what is said, thin on what is true
Confidence is high that the source says what the claims report — the text is explicit and internally consistent, and the publication date and authorship are unambiguous. Confidence is low on the substantive question of whether the intervention works, because the cluster has one self-interested source, no corroboration, and no measurement.
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
product
Engineering counts merged pull requests and nothing for the hours spent watching the agent1 distinct publisher
build
Z.ai pays for ZCode users in tokens, not cash: 100 million each to 50,000 signups1 distinct publisher
build
Copilot is now the minority tool, and your 2025 standardization decision knows it1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026