Skip to content

Build1 publisher3 min readPublished

DigitalOcean's Managed Agents held a running process in memory through a four-minute pause

DigitalOcean Managed Agents brought a paused process back with its in-memory counter intact after more than four minutes, in one developer's independent test. That puts the preview service in a different class from ordinary containers for parking long-lived agents mid-task.

The Engineer · Build desk

Illustration accompanying DigitalOcean's Managed Agents held a running process in memory through a four-minute pause

What happened

  • DigitalOcean's documentation says pausing a Managed Agents session preserves its processes, memory and workspace filesystem, and that resuming restores all three.
  • The agent's own context also came back: after resuming, it recalled a ticket ID it had generated, HARBOUR-7742, from conversation history in eight output tokens without reading the file.
  • A hardware probe found a KVM guest with its own kernel and block device, 2 vCPUs and 3,939 MB of memory, which the tester describes as a microVM.
  • Forking the running sandbox twice took 30.98 seconds and produced three copies of the same process at the same PID, each counting on from its value at the fork.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Choosing where to host a long-running agent now involves a property anyone can test, whether pause keeps RAM, and the write-only counter is a check a team can rerun on any competing sandbox.
  • capability An agent can be run once to a decision point and then continued as several live copies, each starting from the same in-memory state without redoing the earlier work.
  • cost Parking an idle agent and resuming it costs about a fourteenth of the wait for a fresh session, so keeping sessions paused beats rebuilding them on latency.

The obvious version of this test would have proved nothing. The author, who posts on dev.to as remdore, pointed out that writing a file, pausing, resuming and reading it back shows only that a disk survived [15]. "A stopped container loses everything that was not written to a volume, so if that sentence is literally true it is a different kind of thing," the author wrote of the pause claim [3].

So the evidence had to live only in memory. A bash loop held a counter in a shell variable, incremented it once a second and appended it with a timestamp to /tmp/tick.log, a file the process never reads [5]. Had the session been rebuilt from a restored filesystem, the count would have restarted at one. The log shows two adjacent lines instead [7]:

``` 47 08:05:42 48 08:10:10 ```

Both came from PID 590, holding the same variable [7]. The gap between them is 268 seconds, in a loop that sleeps for one [1]. This is good test design, and it is cheap to copy: the whole morning cost about seven cents [4].

Getting the loop to survive took a second attempt. The exec channel into the sandbox kills its children when it closes, and the invocation that lived was this one [6]:

``` setsid nohup /tmp/tick.sh </dev/null >/dev/null 2>&1 & disown ```

Any background helper an agent starts over exec dies the same way unless it is detached [6].

Pause returned in 0.86 seconds and resume in 1.16 [8]. Creating a session took 15.97 seconds from the API call to READY, which the author called "slower than a container and about what a Firecracker-class VM costs you" [11]. The author reported that the paused session consumed nothing during the gap [16]. Sizes run from mars-1vcpu-1gb to mars-16vcpu-32gb [12].

I think the KVM finding explains the result. When the guest has its own kernel and block device [10], the platform can hold the whole machine, and nothing running inside it has to cooperate.

Fork was where the author's assumption failed. The expectation was a disk snapshot and a second sandbox with the same files, which the author said is what checkpoint-and-restore products usually do [17]. The live counter in every copy showed otherwise [14]. The two forks averaged about 15.5 seconds each [2], close to the cost of starting a fresh session [11].

The agent ran on DigitalOcean's own inference endpoint, and one DeepSeek v4 Pro prompt took 7.3 seconds with 15,793 tokens in and 91 out [13]. Inference and sandbox charges land on one account [13].

The evidence is one pause of under five minutes on a service in public preview [1], from a tester who could find nobody else who had checked the claim [18]. The run did not exercise open network connections or pauses measured in hours. The ticker's own timestamps jumped across the pause with no signal to the process [7], so I would expect an agent holding a lease or a token expiry in memory to need a clock check after every resume.

What to watch

  • DigitalOcean stating how long a Managed Agents session can stay paused and what a paused session is billed.
  • A rerun of the counter test with an open network connection held across the pause, or with a pause lasting hours.
  • Whether pause and fork behave the same at the larger sizes, up to mars-16vcpu-32gb.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories