Build1 publisher2 min readPublished
AWS AgentCore Runtime Instances swap per-call cold starts for one always-on GPU host
AWS's AgentCore Runtime Instances put three agents and 18 GB of models on one 24 GB A10G host with a shared, persistent disk. The design cuts cold starts and S3 handoffs on a single host; teams still budget GPU memory themselves and pay for the hours the host sits idle.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Agents pass work by reading and writing a per-session directory on a persistent EBS volume mounted at /mnt/workspace, with the orchestrator handing each one only the session ID.
- The orchestration layer runs one agent at a time by default, and parallel execution has to be configured explicitly with models that fit in GPU memory together.
- If the instance dies, the EBS volume reattaches to a replacement, workflow state in DynamoDB resumes from the last completed agent, and every model has to be reloaded.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction Serial is the default, so on the write-up's own example of three ten-minute GPU jobs the host saves two GPUs and keeps the 30-minute wall time it lists as the problem.
- constraint The S3-free handoff depends on every agent sharing one EBS volume; once a pipeline outgrows one host, EFS, FSx for Lustre or S3 comes back into the data path.
- decision With no per-agent GPU memory cap, one agent's overrun ends in an out-of-memory crash, so teams must budget peak memory across all agents before choosing an instance size.
According to a dev.to summary of AWS's walkthrough, the handoff between agents is a small JSON file [1][12]. When the composer finishes, it writes `.status` into its session directory with four fields: the agent name, a status, the output path and the next agent, here `"next_agent": "arranger"` [12]. The orchestrator reads that file, validates the output and invokes the arranger [12]. Behind it sits a state machine that tracks the session ID, the last completed agent, the next agent, and retry and error state [13]. Workflow state lives in DynamoDB [11]. The agents never call each other. They see paths under `/mnt/workspace/session_abc123/` [4].
I like this design. Every handoff can be inspected after the fact as a file on disk and an item in DynamoDB [12][11].
The cold-start saving is easy to size. The write-up puts a fresh container or Lambda at 5 to 30 seconds per agent call [8]. Three agents means three calls, so a cold pipeline spends 15 to 90 seconds per run just starting up [1]. That range is the write-up's general figure for multi-agent workflows, not a measurement of this pipeline. It transfers only if your own per-call startup lands inside it. A warm host keeps those seconds off every run for as long as the session lasts, and sessions can span multiple days [14].
Memory is where I'd spend the review. With 18 GB of models on a 24 GB card, 6 GB is left over and 75 percent of the A10G is committed before any work starts [2]. The models fit side by side on paper, so this example meets the stated condition for parallel mode [6]. The 8, 6 and 4 GB figures are listed as model sizes [3]. The write-up's sizing rule is the sum of each agent's peak usage [7]. I'd measure those peaks against the 6 GB margin before switching parallel mode on [2].
Cost is the open variable. The write-up does not include an hourly price for the g5.xlarge. The instance bills while agents are idle, and the write-up says Lambda or ECS Fargate is cheaper for infrequent batch jobs [10]. Its example pitches pausing work on Friday and resuming Monday without reprocessing [14]. The host stays warm, and billed, for the whole weekend [10][14]. In my view this design fits a pipeline that moves large intermediate files between agents several times a day on one GPU. For a nightly job, I'd take the write-up's own advice and use Fargate [10].
What to watch
- An hourly price for Runtime Instances on g5.xlarge; it sets how often a pipeline must run before a warm host beats Lambda or Fargate.
- Whether AWS adds per-agent GPU memory limits so one agent's allocation cannot crash the others on a shared card.
- Measured peak memory and wall-time figures from a parallel configuration of the three-agent music pipeline.