Product1 publisher3 min readPublished
Spec-driven development makes teams write down expected behavior and acceptance criteria before an AI agent edits code
System Design One's newsletter says AI agents on big codebases need written specs, citing a change that passed tests but emailed an opted-out user. Most of its sample spec just records what the old system already does, and a team can write that document before buying any agent platform.
The Product Desk · Product desk

What happened
- The newsletter says the trouble grows with codebase size and agent workload, naming three problems: understanding the existing system, building the right change, and trusting autonomous work.
- Its sample spec limits the move to task-assignment and task-comment notifications, leaves billing notifications unchanged, and covers edge cases such as failed delivery.
- It describes three roles for a spec: written before implementation, kept connected to the code afterward, or treated as the source from which plans, tasks and tests are generated.
- Its case study is Blitzy, an enterprise platform that combines a team's requirements with codebase context across repositories to produce what it calls an Agent Action Plan.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint A green test suite is weak sign-off for any agent migration that must keep settings users chose under the old system, since it only checks what someone thought to test.
- decision Teams handing an agent a migration now have to assign someone who knows the old system to write down what it does, and that has to happen before the agent starts.
- exposure As agents get more autonomy, one vague instruction can reach more files and tools, so access limits and human approval points have to be part of the rollout plan.
The user in this story did one thing: at some point they switched email notifications off. After the migration, a colleague assigns them a task and the email arrives anyway, and the test run stays green [1]. The newsletter wrote, "The code passes tests, but it does NOT behave the way you want." [14]
Handing an agent a migration with a short prompt assumes it will learn the old behavior by reading the code. The newsletter disagrees. It says a prompt describes what you want and does not explain how the existing system works, including feature flags and data models spread across repos [4]. Without that context, it says, the agent makes a locally correct change that breaks something elsewhere [5].
The obvious patch is a better prompt, with lines like "Preserve existing notification preferences" [2]. The newsletter's objection is scale. Adding more instructions makes the request hard to manage, and a vague instruction can quietly become a wrong implementation repeated across many files or tasks [6]. A spec, as the newsletter defines it, sets expected behavior, scope, constraints and success criteria before the agent starts [8].
The product being pitched is Blitzy's platform [11]. The work being described is a document. Most of the sample spec records what has to survive the move: existing preferences, delivery timing and the current email service [16]. The line that would have stopped the stray email is under expected behavior. It says to check the user's preferences before sending, and to queue an email through the new path only if email is enabled [16]. The acceptance criteria then require existing preferences to "continue to work correctly" [16].
The argument rests on a scenario that begins with "Imagine" [15], and it comes with no measured rate of regressions like this one.
I think the recommendation stands without the platform. Before an agent touches behavior that users configured, write down what the old system does for them and turn the "unchanged" lines into acceptance tests. That costs time up front from whoever knows the old system. There is also upkeep later if the spec stays connected to the code to guide future changes, the role the newsletter calls spec-anchored [9].
Two axes sort the work. One is whether the change must preserve something a user set, such as an email opt-out, or only adds something new. The other is whether the behavior involved sits in code the agent can see or is spread across repos and feature flags [4]. A contained change that only adds behavior can reasonably ship on a clear prompt and a green test run. An additive change that spans systems needs the scope and dependency lines of a spec. A contained change that must preserve user settings needs those settings written as tests against today's behavior before the agent starts. The notification migration is the fourth case: preserved behavior spread across systems. It gets the full spec, edge cases included, and a person approves the agent's plan before it runs.
What to watch
- Whether the newsletter or Blitzy publishes measured results, such as regression rates on agent-built migrations with and without a written spec.
- Whether agent coding tools start shipping spec templates with a required section for behavior that must stay unchanged.