Skip to content

Build1 publisher3 min readPublished

Context rot at 15 iterations: two toolkits that move the spec into Git

A dev.to comparison of OpenSpec and GitHub Spec Kit argues chat degradation is structural. The interesting part is not the diagnosis but how differently the two tools file the cure.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Context rot at 15 iterations: two toolkits that move the spec into Git
Generated illustration

What happened

  • The dev.to post describes "vibe coding" as prompting an LLM in an open chat window, hitting apply, and tweaking code until the test suite or browser stops throwing errors.
  • The post asserts that as chat conversations grow beyond 15-20 iterations, the LLM loses track of earlier decisions and starts reverting fixes or introducing regressions.
  • The post asserts that when code is generated directly from conversational prompts, the reasoning vanishes once the chat window closes ("lost architectural intent").
  • The post asserts that reviewing a 1,500-line diff generated across multiple chat sessions is exhausting because the requirements were never codified in Git.
  • The post breaks spec-driven development into phases: Constitution/Rules (global invariants such as tech stack, security rules, dependency budgets, coding style), Intent (user stories, business constraints, Given/When/Then acceptance criteria), Architecture (system design, data schemas, API contracts, component boundaries), and Execution (a dependency-ordered checklist of atomic implementation steps).

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A walkthrough published on dev.to compares two open-source spec-driven development toolkits, Fission-AI's OpenSpec and GitHub's Spec Kit, and hangs its argument on a claim worth separating from the tooling [6][8]. According to the post, chat conversations degrade past 15 to 20 iterations, with the model losing track of earlier decisions, reverting fixes and introducing regressions [2]. If that failure is a property of the medium rather than of the operator, no amount of prompt craft closes it, and the specification has to be written down somewhere a reviewer can see it.

The post lists two consequences that matter more than the first. Architectural intent disappears when the chat window closes, because the reasoning was never recorded anywhere but the transcript [3]. And a 1,500-line diff assembled across several sessions is unreviewable, since the requirements it was supposed to satisfy were never committed to Git [4]. That is the operational cost: not bad code, but code nobody can approve or reject on the evidence.

The proposed mechanism is a layered, fixed context. The post argues that by prompt 15 a model spends 80 percent of its attention budget parsing its own prior mistakes [10], and offers no measurement for that figure, so treat it as illustration rather than finding. The alternative it sketches is more concrete: each task executes in a clean context window [17] loaded with a constitution at roughly 500 tokens, a spec at 800, a plan at 1,000 and about 400 tokens of active task scope [11]. That totals around 2,700 tokens of standing context per step [18], which is the actual claim being made. A fixed input is cheaper and more predictable than an accumulating one, whatever the attention-budget story turns out to be.

Where the two implementations diverge is in what they ask you to write first. OpenSpec targets brownfield repositories and multi-agent work, and does not require documenting a legacy codebase up front [6]. You propose an atomic change, write delta specs describing only what shifts relative to the current system, and once tests pass the change is synced into permanent specs and archived [7]. Spec Kit, GitHub's toolkit driven by the specify-cli Python tool [8], starts at the other end: a constitution.md defining inviolable rules for architectural patterns, linting, test coverage and security boundaries before any feature is specified [9]. One is designed for repositories that already exist and cannot be re-documented; the other is designed to constrain a project before it accumulates habits.

The practice notes are the least glamorous and probably the most portable. State non-functional constraints explicitly rather than assuming the model infers them, as in a bundle size ceiling of 5KB gzipped [12]. Write acceptance criteria as Given/When/Then scenarios, such as an expired token returning HTTP 401 with code TOKEN_EXPIRED [13]. Decompose tasks to single-file or single-function units instead of "Implement auth" [15]. Get human approval on the spec before code, on the grounds that a 30-line Markdown file takes 30 seconds to correct [16].

What to watch: whether either project publishes an actual measurement behind the 15 to 20 iteration threshold [2], and whether delta specs stay in sync with the code they describe once a team stops archiving them diligently [7].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories