Skip to content

Build1 publisher3 min readPublished

LLM-dependent agents flipped HivePlane's certification verdicts between runs

Debashish Ghosal got HivePlane's agent certification to repeat across three runs by swapping in deterministic agents and a mocked judge. The stable verdict covers the control plane's plumbing, and whether an LLM-driven agent answers correctly is now outside the certificate.

The Engineer · Build desk

Illustration accompanying LLM-dependent agents flipped HivePlane's certification verdicts between runs

What happened

  • Three agents the author had already built, a repo guardian, a release-notes drafter and an incident commander, failed on day one from a missing dependency, a missing package layout and the wrong SDK.
  • A live-stack run found api.pagerduty.com missing from the sandbox egress allowlist, so a destructive acknowledge call was denied and the run died.
  • A paused LangGraph run survived a restart of the API container with its 8-event log intact and resumed to completed from a durable checkpoint.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Anyone building an agent certification gate has to choose whether the certificate covers the plane or the agent's judgment, because this gate only reproduced once the model's opinion was mocked out.
  • constraint On six- and four-task corpora a 0.90 threshold demands a perfect score, so any LLM left in the verdict turns one misclassification into a failed certificate.
  • cost Agents written against one repo's dependencies and SDK need packaging work before a fleet platform can run them, and in the author's account that work falls on the agent's owner.
  • exposure Egress gaps like the missing PagerDuty host stay hidden while unit tests mock hosts, and they surface only when a destructive call hits a live sandbox.

The flip started at a classifier. The agents drive human-in-the-loop flows whose outputs depend on an LLM [5]. The real local model labelled a bug-fix task as a changelog and failed certification nondeterministically across runs [6]. The pass thresholds never moved. "My thresholds were deterministic; my signal wasn't," Debashish Ghosal wrote [7]. He was blunter from the platform's side: "A certification result that changes run to run isn't a certification. It's a coin flip with a signature on it." [8]

His fix was to decide what the field test measured. Ghosal spent six weeks building HivePlane [1]. He describes it as a loop: register, certify, gate admission, enforce budget and policy, pause and resume, deliver and audit. In his framing the agent is the workload and the plane is the product [9]. So the Tier 1 subjects became deterministic agents wired through thin shims, with a mock knowledge base, mock tools and a mock judge [10]. The same seed always produces the same category of run [10]. Identity binding, the provider seam and the model-swap gate still execute during certification. "What is mocked is the judge's opinion, not the plumbing," he wrote [11].

The S1 suite then passed identically in three consecutive stack runs [12]. Support-agent certified at 6/6 with a p95 of 67 ms, and eval-judge at 4/4 with a p95 of 126 ms, both at a 0.90 threshold with signed Ed25519 attestations [13]. I think mocking the judge is the right call for a control-plane test. A gate that reads a sampled model output cannot give the same answer twice on demand. The cost is scope. The attestations cover the plane's loop. They do not tell you whether an LLM-driven agent will label a bug fix correctly next Tuesday. For the 6/6 to transfer to your own agent, its decisions would need to be seeded the same way, or you would have to accept a mock judge inside the certificate.

Corpus size matters as much as the seed. At six tasks, one miss gives 5/6, or 0.83. At four tasks, one miss gives 0.75 [14]. Both fall under 0.90, so at these sizes the threshold requires a perfect score [14]. With the LLM still in the verdict, one bad classification would fail the whole certificate. The latency figures are thin for the same reason: a p95 over six tasks falls between the fifth and sixth ordered samples, next to the single slowest task [15].

The portability failure came first, and Ghosal wrote that none of it was the control plane's fault [2]. His own three agents failed in the new environment for three different reasons: a missing dependency, a missing package layout and the wrong SDK [3]. He wrote that an agent working in its own repo does not amount to an interface [4]. He also wrote: "If your agent can't run outside its birth repo, it can't be operated by a fleet." [20]

The live stack found faults the unit tests could not. An escalation task issued a destructive pagerduty.acknowledge call, and api.pagerduty.com was missing from the manifest's sandbox.egress.allow list [16]. The call was denied. The run died with expected status='escalated', got None, because the unit tests had mocked the host [16]. Two other scenarios kept being aborted mid-run because they wrote no evidence until after the final poll, so each abort erased the diagnosis [17]. The fix writes submitted.json, paused.json, approvals.json and resumed.json before each boundary [17].

The pause-and-restart case had worked the whole time. A paused LangGraph run survived docker compose restart api with its 8-event log intact, re-attached from a durable checkpoint on the volume, and resumed to completed [19]. Anyone scripting it should expect POST /runs/{id}/resume to return 409 when the approval's automatic re-dispatch has already advanced the run. Success means the run reaching completed [18].

What to watch

  • Whether HivePlane adds a certification tier for LLM-driven agents that runs each task repeatedly and passes on a rate, so model drift is measured without flipping the gate.
  • Whether the repo guardian, release-notes drafter and incident commander are repackaged to run drop-in and put back through the certified loop.
  • Certification corpora larger than six and four tasks, where a 0.90 threshold would tolerate a failure and stop demanding a perfect score.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories