Build1 publisher3 min readPublished
Relipa's coding-agent pipeline sets the tests before the agent writes any code
Relipa's AI development workflow wraps a coding agent in six stages and fixes the tests before any code is written. Most of its published detail covers the workspace plumbing that runs before a model is ever called.
The Engineer · Build desk

What happened
- If the worktree for any one repository fails, the system removes every worktree it has just created.
- Branch slugs are NFD-normalized with the letter đ mapped by hand, capped at 40 characters and suffixed with a four-hex workspace ID.
- The available post text breaks off inside the Spec stage, before its sections on the capped coding loop and on where the workflow leaks.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams copying the test-first order must decide who writes the tests and whether the agent may edit them, because independence is the stated reason for the order.
- cost Before the agent contributes anything, a team has to build idempotent, transactional workspace handling across every repository a task touches.
- constraint Sequential setup holds the agent until the environment is ready; parallel mode buys a faster start at the risk of test failures caused by setup still running.
The post's author, describing work at Relipa, wrote that "the interesting part isn't getting an AI agent to write code. It's designing what happens before and after the code is written" [1]. The same post describes the job this way: "The developer's role shifts from writing every line of code toward deciding what the agent should do, what it should not do, and how we can prove that the result is correct" [2].
The last clause of that quote explains the stage order. The team defines what should be tested before the agent implements anything, so the tests can act as an independent reference when the code is judged later [4]. Review then sends inline comments back to the Code stage [5]. I think the order is right, with one condition. The tests stay independent only if they are frozen before the agent starts and the agent cannot edit them to turn a failing run green.
The best-documented stage calls no model. A card dragged into an auto-run column, or a task created by hand, triggers workspace allocation first [13]. The system creates a separate git worktree and branch for every repository in the task [6]. Two rules make that safe to run twice. Each issue has at most one workspace, so dragging a card back resumes the existing worktree and branch. Without that rule, the author wrote, a mis-drag can leave an orphan branch and no clear answer about which branch is active [7]. Creation is also all or nothing: if any repository's worktree fails, every worktree just created is removed [8]. Those are the properties you want from any job runner. It is idempotent on the issue and transactional across repos. Multi-repo tasks are what force the second rule [6][8]. An agent dropped into two of three repos would be working against a half-built environment.
Branch naming gets the most careful handling. Slugs are normalized to NFD, stripped of combining marks, capped at 40 characters, cut back to a word boundary and suffixed with the workspace's four-hex-character identifier [9]. The letter đ is mapped separately because it has no decomposed form [9]. That is good Unicode work. Without proper normalization, "Sửa lỗi đăng nhập" can turn into "s-a-l-i-ng-nh-p", which the author called "technically usable but almost unreadable" [10]. Git will take that name without complaint. The four-hex suffix allows 65,536 distinct values [1].
At the workspace root the agent gets copied environment variables and two instruction files, CLAUDE.md and AGENTS.md. They hold @<repo>/<file> references to each repository's conventions [11]. A repo-level setup script runs next. In the default sequential mode the agent waits for it to finish; in parallel mode the agent starts immediately [12]. Sequential is the default I would pick. An agent that runs tests while dependencies are still installing will report failures that have nothing to do with its code.
Spec and Plan run together. Spec turns the ticket into user stories with P1, P2 or P3 priority and acceptance criteria for each story, and Plan covers architecture and data model [14]. The available text of the post ends partway through that section. Its contents list later sections titled "Code: a loop with a ceiling" and "Where it still leaks" [15]. The retry limit on the coding loop and the failure cases the author found are therefore not in this record.
What to watch
- Whether the full post states the retry ceiling on the Code loop and what the pipeline does when an agent reaches it.
- Whether Stage 4 tests are written by a human or by the agent, and whether the agent can modify them during Code.
- What the post's 'Where it still leaks' section names as the failure points across the six stages.