Skip to content

Build1 publisher3 min readPublished

AWS lifts the eight-hour cap on Bedrock agents by putting sessions on your own EC2

AgentCore runtime instances stretch agent sessions from eight hours to fourteen days. The architecture problem becomes a configuration choice, and a utilisation bet.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying AWS lifts the eight-hour cap on Bedrock agents by putting sessions on your own EC2
Generated illustration

What happened

  • AWS introduced runtime instances in Bedrock AgentCore, a second compute option that runs agents on managed Amazon EC2 in a customer account while keeping the existing AgentCore APIs, identity and observability model.
  • AgentCore launched with serverless microVM-based sessions subject to an eight-hour ceiling.
  • Principal developer advocate Sebastien Stormacq says runtime instances give agents persistent EC2-backed sessions that can run for up to fourteen days, with shared file systems, GPU accelerated instance types, and support for Python and container images.
  • Multiple agents can be deployed into a single runtime and collaborate on the same host through a shared session directory, rather than calling each other's APIs for every handoff.
  • Stormacq: "your agents can call each other as tools within a shared session, iterating autonomously until the job is done."

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Amazon Web Services has added runtime instances to Bedrock AgentCore, a second compute option that runs agents on managed EC2 inside a customer account while keeping the existing AgentCore APIs, identity controls and observability model [1]. AgentCore launched on serverless microVM sessions with an eight-hour ceiling [2]; runtime instances raise the session limit to fourteen days [3], which is 336 hours, or 42 times the old wall [16].

That number matters less than what it removes. Anything that needed to survive longer than a working day previously had to be decomposed into resumable chunks with externalised state, or moved onto a separate EC2 fleet the team maintained itself [11]. Now the same work is a session time-to-live setting. Runtime instances also bring shared file systems, GPU-accelerated instance types, and both Python and container image packaging [3]. Several agents can be deployed into one runtime and collaborate on the same host through a shared session directory rather than calling each other's APIs on every handoff [4]. Principal developer advocate Sebastien Stormacq writes that agents "can call each other as tools within a shared session, iterating autonomously until the job is done" [5], and that CrewAI, LangGraph, LlamaIndex and Strands can be brought across without changing packaging, using an `@app.entrypoint` decorator plus a zip file or container image [6].

Underneath sits a new primitive, the capacity provider, which defines allowed instance families, operating system, networking and storage, and acts as the contract between agents and the EC2 capacity AgentCore provisions, patches and scales [8]. Teams set minimum and maximum instance counts and a target utilisation, then attach one or more agent runtimes with a session TTL of up to fourteen days, without hand-managing Auto Scaling groups, launch templates or AMI pipelines [9]. According to the Enkompass engineering guide, this is the right home for workloads that need more than eight hours of continuous runtime, GPUs or large memory, or tight co-location; the microVM runtime stays the default for short, bursty, request-response agents because it starts quickly, isolates sessions, and bills per second on actual CPU and peak memory up to eight hours [10][11]. AWS presents the two as complementary and suggests mixed topologies, with an orchestrator on microVMs dispatching long-running work to instance-backed workers [7].

The bill changes shape. A cost analysis from eCorpIT notes that runtime instances are billed at standard EC2 rates for the chosen instance type plus a management fee, while microVMs charge per vCPU-hour and per GB-hour with no management fee and come out cheaper for agents with low sustained CPU and long idle gaps [12]. eCorpIT puts the break-even at roughly 24 percent sustained CPU utilisation before Savings Plan discounts [13], which means an agent idling below that line costs more on an instance than on a microVM [17]. Savings Plans and Reserved Instances reach the compute portion but not the management fee [14], so the levers that actually move the number are packing multiple agents onto one host and stopping sessions the moment work finishes [15].

There is an isolation trade here too. MicroVM sessions isolate from each other by construction [10]; the co-location pattern deliberately puts multiple agents on one host with a shared working directory [4], which makes the blast radius of a misbehaving agent a design question rather than a platform guarantee.

Watch whether teams actually stop sessions, or leave fourteen-day TTLs running at single-digit utilisation [13][15]. Watch how much of a fleet ends up on each model once both are in production [11].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories