Build1 publisherNot yet confirmed elsewhere3 min readPublished
Mainbrella reserves capacity before boot to keep agent retries from starting a second machine
Mainbrella's agent sandbox keeps idempotency keys for 24 hours so a retry after a lost reply finds the first machine instead of billing a second. Its published test covers admission under contention, so how those retry paths behave at production scale is still unmeasured.
The Engineer · Build desk

What happened
- Andrew Arrow's October 7th engineering post describes the failure behind the design: a cloud machine starts, the reply is lost, and the agent retries by creating another.
- Mainbrella sells Linux workspaces that AI agents create and control through an API, with its Builder plan listed at $5 a month.
- If the original machine has stopped or its slot was reused, the API returns creation_no_longer_running and the client must send a new key to get a replacement.
- When a runtime restarts before a managed job finishes, Mainbrella marks the job interrupted and does not rerun the command on its own.
- A repository test fires six simultaneous creates at a five-slot Builder account; five machines reach boot and the sixth request gets a conflict.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Agent clients need an explicit branch for creation_no_longer_running, where someone decides whether a replacement machine is worth a new key and a new start.
- cost Callers absorb lost output when a job is interrupted, and long-running jobs need resume or rerun logic written on the client side.
- exposure Anyone running agent fleets on this design at volume is relying on reconciliation and fencing paths that the published test does not exercise.
Ordering is the part I would copy. When a create call boots first and updates capacity afterwards, two simultaneous requests can both see the last free slot and both start machines [5]. Mainbrella's account-level coordinator, a Cloudflare Durable Object, reserves the slot and the monthly start and compute allowance before anything is dispatched to boot [4]. The idempotency receipt lands in the same atomic storage write as the reservation [8]. So no slot is ever claimed without the key that claimed it being recorded too [8].
The lock is short. It is released once the reservation is durable, and the runtime then boots on its own [6]. Admission is serialized per account while boots overlap [4][6].
Duplicate protection depends on the caller. The receipt is found by its Idempotency-Key, so a framework that mints a fresh key on every attempt is making a new request each time and can claim a new slot [18]. A retry that finds its reservation still pending waits through a 90-second reconciliation window, and then the coordinator checks what actually exists at the runtime [9]. The 24-hour key retention is 960 times that window [19]. An agent can come back hours after a dropped reply and still be pointed at its first attempt [8][10].
Reused slots get the most careful handling. Each admitted start takes an increasing reservation number, and the runtime records cancellation fences against those numbers [11]. A delayed boot for an older reservation is rejected, and a late cleanup aimed at the old machine cannot stop its replacement [11]. Commands and file requests also carry a machine generation, so a slot name alone cannot authorize work against a later lifetime of that slot [12]. These are fencing tokens, applied at both the boot and the command layer [11][12]. It is careful work from what the company's homepage describes as a sole proprietorship in Culver City, California [3].
Job execution follows the same rule. Managed jobs have their own idempotency keys and saved execution records, and a client that drops its connection can resume reading output from a stored sequence cursor [13]. Runtimewire's account states the cost of declining to rerun: a script writing a CSV can lose unfinished work, while replaying an operation with side effects could be worse [15]. I think that is the right default for agent workloads. The agent that issued a command can decide to run it again. The platform cannot tell whether a half-finished command already took effect.
Readiness is a lower bar than the word suggests. The check runs uname -a, which establishes that the kernel can print its own name [16]. It confirms the guest can execute a command, and it does not show whether the agent's application ran correctly [16].
The published test checks one race. Holding readiness behind a gate takes boot time out of the result, and runtimewire says the test does not measure production startup speed [7]. For the result to say anything about reliability, a test would have to drop replies during creation, restart a runtime mid-job, and deliver a late command to a reused slot. Those are the cases the reconciliation window, the interrupted-job state and the fences were built for [9][14][11]. Runtimewire's own summary is that the test checks exclusive admission and parallel provisioning, not reliability at production scale [17].
What to watch
- Whether Mainbrella publishes tests that drop creation replies or restart runtimes mid-job, exercising the 90-second reconciliation and interrupted-job paths.
- How agent SDKs and frameworks that call the API generate Idempotency-Key values across retries.
- Production startup and failure-rate figures, which the current repository test does not measure.