Skip to content

Build2 publishers3 min readPublished

AWS's ambient-agent sample exits Lambda at each human checkpoint and resumes from DynamoDB

AWS published a Bedrock AgentCore sample for agents that wake on S3 and scheduled events, capping each agent turn at Lambda's 15-minute timeout. Reviewers can take hours, so the sample stops the agent at every human checkpoint and restarts it from saved state.

The Engineer · Build desk

Illustration accompanying AWS's ambient-agent sample exits Lambda at each human checkpoint and resumes from DynamoDB

What happened

  • Events from S3 notifications, EventBridge schedules, CloudWatch alarms and SNS topics all land in one SQS queue, which triggers a Lambda function that runs the agent.
  • The ask_human tool is the only one that pauses execution, while reading S3, querying a database and sending email all run synchronously inside the same invocation.
  • At a pause, the handler saves the agent's reasoning chain, its pending decision and an execution ID to DynamoDB, then exits.
  • SQS cannot route by message content, and the dev.to writeup says AWS's post does not specify how the Lambda decides which agent workflow a message belongs to.
  • AWS says one ask_human tool plus a canonical response envelope is enough to handle the full range of human-in-the-loop interactions.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Before writing any agent logic, a team running the sample takes on resources in nine AWS services ahead of Bedrock, from IAM and Cognito to CloudFront and ECR.
  • decision Each adopting team has to write and maintain its own dispatcher from message body to workflow, and every new event source means changing that one function.
  • exposure Synchronous tools have already run when the agent reaches ask_human, so a rejection cannot recall an email the agent sent earlier in the same turn.

The design follows from one constraint: where the waiting happens. A chat agent runs inside a request-response cycle. It keeps state in memory or a session store, and the conversation ends when the user closes the window [13]. An ambient agent's reviewer may answer hours later [14]. The reference caps each agent turn at Lambda's 15-minute timeout [2]. A wait of hours does not fit in a 15-minute invocation, so the agent has to stop running while the human decides [2].

AWS's answer is to exit early. When the agent calls ask_human, the tool returns a special status code. The handler catches it, serializes the agent's state, writes it to DynamoDB and exits [5]. Approval comes back as a second SQS message that triggers the function again [7]. Nothing runs while the reviewer is at lunch [5]. I think this is good engineering. Every pause goes through one tool, so exactly one code path writes a paused agent to the table [4].

Resume is where it can break. The dev.to writeup says the agent picks up from the exact tool call that triggered the pause [7]. AWS's own wording is that it resumes from where it left off [1]. That only works if the resumed agent treats the saved tool outputs as finished. Tools such as sending email run synchronously inside the invocation [4], so any that ran earlier in the turn have completed before the pause. An email sent ahead of the checkpoint is already out when the reviewer says no.

The routing is code the team writes. The writeup offers an illustrative handler and labels it a typical pattern. It checks for eventSource 'aws:s3', source 'aws.events' and an AlarmName field, then falls back to 'default-agent' [9]. Every new event source is another branch in that one function.

Concurrency comes from Lambda. Each invocation takes one SQS message, and workflows with different execution IDs run in parallel [10]. When a resume shares an execution ID, DynamoDB optimistic locking prevents the race, according to the writeup [11]. The Jobs page finds pending work by polling DynamoDB for rows marked WAITING_FOR_HUMAN [12].

The 15-minute figure is a claim about AWS's workload. AgentCore Runtime offers container hosting for long-running workloads with session isolation [17]. This reference still stops each turn at Lambda's limit, and AWS calls that "more than enough headroom in practice" [2]. The practice AWS describes is a document landing in S3 and an agent analyzing it before asking for approval [20]. For the claim to carry over, a team's slowest single turn, model calls plus synchronous tools, has to finish inside 15 minutes.

The setup list is long. The prerequisites ask for permission to create IAM roles, Lambda functions, DynamoDB tables, S3 buckets, SQS queues, API Gateway APIs, CloudFront distributions, Cognito user pools and ECR repositories, plus Bedrock resources [18]. Nine AWS services appear in that list ahead of Bedrock [1].

The writeup's author put the conclusion bluntly: "This is not a chatbot with extra triggers. It is a different control flow." [19] The sources support that. The human step covers more than approval, though. AWS says the agent interrupts a person when it needs clarification, approval or review [16].

What to watch

  • Whether AWS publishes the schema of its canonical response envelope and the event-routing code in the reference sample repository.
  • Whether a later version of the reference runs agent turns on AgentCore Runtime directly and drops the Lambda 15-minute cap per turn.
  • How the sample records or compensates for side-effecting tool calls when a reviewer rejects the decision at ask_human.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories