Skip to content

Build1 publisher2 min readPublished

Agent work needs separate fields for acceptance, merge and deployment, a dev.to tutorial argues

Splitting an AI agent's 'done' into three status fields stops a green CI run from standing in for a reviewer's acceptance, a dev.to tutorial argues. That costs 16 state values to track and a permission check on each transition.

The Engineer · Build desk

Illustration accompanying Agent work needs separate fields for acceptance, merge and deployment, a dev.to tutorial argues
Generated illustration

What happened

  • The work tracker owns intent and review, the repository and release system own integration and deployment, and the post says not to force them into one status field.
  • Under the ordinary review path a runner cannot accept its own delivery, and a successful deployment cannot invent a missing delivery report.
  • A retried delivery carrying the same operationId returns the original result, so a client timeout cannot create a second delivery record.
  • An expected ticket revision stops a stale browser tab or a delayed agent process from delivering against an attempt that was reopened or reassigned.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure If a team grants the reviewer role to an agent identity, one agent can accept another's work and the record will still show independent acceptance.
  • cost Every client that delivers work, agents included, has to send an operationId and an expected revision, and the server has to keep operation results to answer retries.
  • decision Teams adopting the model have to choose which system writes each field, and which task types get not_applicable for merge and deployment.

"Done" in this design is three fields. The work state runs draft, open, claimed, delivered, accepted, and ends there [3]. Integration has six values: not_applicable, unreported, branch_reported, pull_request_open, merged and blocked [4]. Deployment has five: not_applicable, not_deployed, deploying, deployed and failed [5]. A non-code task can be accepted with neither of the other two fields applying [8]. "An accepted task can still be unmerged. A merged change can still be undeployed," the post says [7].

Permissions attach to named roles. The post defines six: requester, task_owner, runner, reviewer, repository_maintainer and release_owner [9]. One person may hold several, but each check should name the role being exercised [10]. The post gives the reason: an administrator override should not look like an ordinary peer review [10].

The accept guard tests three things [13]. The ticket must be in delivered. The actor must hold the reviewer role. The actor's ID must differ from the runner's. If that last test fails, the error reads "Independent acceptance requires another person" [13]. The error message promises a person, while the code only compares two ID strings. A second agent with its own ID and the reviewer role passes all three tests [3]. In the prose, the different-person rule applies to higher-risk work, and where self-verification is allowed it has to be recorded explicitly [11]. The function as shown applies the rule to every acceptance [13].

The delivery path gets an ordering detail right. Its first step returns any stored result for the operationId. Only after that does it load the ticket, compare the revision and check state and runner [19]. Take a client that times out after the server has stored its delivery. The ticket now sits in delivered at a new revision. If the retry hit the state check first, it would throw "Only claimed work can be delivered" [14] and the client would see a failure for a write that had already landed [2]. Because the operationId lookup comes first, the client gets the original result back. The fourth step inserts the delivery and the state transition atomically [19].

According to the post, all of these checks belong on the server. "Hiding a button in the interface is useful feedback, not authorization," it says [15]. It also limits what a delivery claims: "this attempt is ready for judgment," which does not mean the work is correct [16].

The available text stops at the review-failure section, which opens with "Review failure is not one situation" [21]. The only guarded code it shows is for delivery and acceptance [13][14].

What to watch

  • Whether the rest of the post shows guarded transitions for merge and deployment, and which roles it requires to make them.
  • How the review-failure section separates sending work back to the same runner from reopening it for a new attempt, which is where attempt IDs and expected revisions get tested.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories