Category Archives: AgentCore

Agents Aren't Web Requests — Why AWS Rebuilt the AgentCore Runtime

Years ago I sat with a customer who needed to orchestrate multiple agents inside their software products — agents that coordinate, hand work to each other, and stay alive across a workflow, not one model answering one call. We wrote a PRFAQ together, took it into an EBC, and a VP agreed to build it. The document was right. It was also complex — it described, in detail, nearly every capability you can find in AgentCore today. The service team then rewrote it, more than once, not to correct it but to break a correct and complex vision into something they could actually execute and build. The primitives got names and edges they did not have in our draft. But the shape of the thing never changed, because we had the mental model right from the start: these are not web requests. They are sessions.

That is the part I am proud of, and it is the only part that matters here: the vision was correct and complete, and the one idea holding all of it together was that an agent is a session, not a request.

Here is the mental model that quietly breaks the moment you ship an agent to production: “an agent is just a web request that happens to call an LLM.” You wire up an HTTP handler, it fires off a prompt, you stream tokens back, the function returns, the compute is reclaimed. Clean. Stateless. Familiar.

It is also wrong in ways that will cost you. An agent is not a request. It is a process — one that reasons over many turns, accumulates intermediate results, calls tools that mutate state, occasionally reaches for a GPU, and sometimes hands work to other agents. The shape of that workload has almost nothing in common with the shape of an HTTP handler. Amazon’s answer, Amazon Bedrock AgentCore Runtime, is best understood not as “serverless with a bigger timeout” but as a deliberate rethinking of what compute for agents should look like.

What breaks when you run an agent like an HTTP handler

Serverless web infrastructure makes three assumptions that are load bearing for request response traffic and actively hostile to agents.

It assumes statelessness. An HTTP handler is supposed to carry nothing between invocations; that is what lets the platform scale it horizontally without thinking. But an agent’s whole value is the state it carries — conversation history, a scratchpad of intermediate reasoning, the output of a tool it called two steps ago, a half built file on disk. If every invocation lands on a fresh execution environment with an empty filesystem, you are forced to serialize and rehydrate the entire context on every single turn. That is slow, lossy, and fragile.

It assumes short duration. Request response compute is tuned for work that completes in milliseconds to seconds, with hard timeouts measured in minutes. Agents routinely run for minutes to hours, and some genuinely need to run continuously for days — a transformation job grinding through a corpus, an automation loop that pauses and resumes, a research agent that keeps working while a human is asleep.

It assumes weak isolation is fine. When requests are stateless and ephemeral, co tenancy is mostly an efficiency detail. But agents run privileged, nondeterministic code on a user’s behalf — they get shell access, read and write files, and invoke tools with real credentials. The isolation boundary stops being a performance concern and becomes a security boundary. You cannot have user A’s agent able to observe anything left behind by user B’s.

How AgentCore answers it: isolated sessions, not shared handlers

AgentCore Runtime’s primitive is the session, not the request. Each user session runs in its own dedicated microVM with isolated compute, memory, and filesystem, plus shell access. One user’s agent cannot reach another user’s data, and when the session ends the entire microVM is torn down and memory is sanitized — no cross session residue. That is a deterministic isolation boundary wrapped around a deliberately nondeterministic workload, which is exactly the property enterprises need.

Inside that boundary, the model does the thing HTTP handlers refuse to do: it lets you safely reuse context across invocations. You generate a session ID, pass the same ID on every related call, and each invocation builds on the environment the previous one left behind — same filesystem, same in memory state, same tool context. Multi turn conversations and multi step workflows stop being a serialization problem and become what they always should have been: successive calls into a living environment.

The honest part of the design is that this state is ephemeral by default. In memory and on disk data lives for the session’s lifetime and no longer. For anything that must outlive the session — learned preferences, durable conversation history, workspace files — you pair the runtime with AgentCore Memory for long term recall and Amazon EBS for persistent volumes. The runtime handles scaling, session management, isolation, patching, and observability so you spend your time on agent logic, not plumbing.

Two compute shapes, one set of APIs

The sharpest architectural decision is that AgentCore exposes two compute shapes under the same runtime APIs, so you match the shape of the compute to the shape of the traffic.

Serverless microVMs are the default: fast cold starts, scale to zero, sessions up to 8 hours. This is the right home for request response, spiky, or latency sensitive agents — the chat assistant, the API driven tool, the thing that needs to start instantly and cost nothing while idle.

Runtime instances, generally available since August 6 2026, are AWS managed EC2 running in your own account, defined by a reusable capacity provider (operating system, allowed instance types, networking, storage, IAM roles). Sessions here run up to 14 days, support GPU instance families with drivers provisioned for you, and can stop and restart to save cost — hibernate Monday night, resume Wednesday morning with volumes reattached and data intact. This is the home for long running, continuous, GPU bound, or multi agent work.

Because the instances live in your account, your data and your account controls stay put, and you can apply existing Savings Plans, Reserved Instances, and ODCRs. AgentCore still owns the lifecycle — provisioning, patching, scaling, teardown — so you get EC2 economics without EC2 babysitting.

Multiple agents on one host

Runtime instances change the unit of collaboration. In the microVM model, one runtime hosts one agent. On an instance, a single session can host many agents that share a filesystem. AWS’s own demo makes the point: a writer agent generates code into a shared session directory, and a reviewer agent reads that same file and critiques it — no message passing, no inter agent API calls, no data transfer. They collaborate through the filesystem.

That composes cleanly with the two shapes. A lightweight orchestrator on a microVM can take fast, spiky, API driven traffic and dispatch heavier work to specialized worker agents running on instances — code compilation, security scanning, GUI automation — that need persistent state and direct OS access. And none of this locks you into a framework: bring CrewAI, LangGraph, LlamaIndex, or Strands, and any model. Packaging is minimal — an @app.entrypoint decorator plus a zip or a container image.

The caveat worth saying out loud

Here is where I will push back on the easy narrative. A 14 day session is not a durable workflow. If your process waits days for a human approval, or moves money, or must survive infrastructure failure with transactional guarantees, a single long lived session is a fragile place to park that state. Runtime instances raise the ceiling for continuous work; they do not turn one invocation into a saga. For orchestration that must be durable and recoverable, reach for a real orchestrator — AWS Step Functions over a transactional store — and let the agent be a step inside it, not the system of record.

Two more things to keep honest. Instances do not scale to zero; an idle instance still bills until you stop the session. And the whole point of the two shapes is that neither is universally correct. Match the compute shape to the traffic shape or you will overpay for idle capacity or starve a long job of the persistence it needs.

Takeaways

  1. Model agents as sessions, not requests. Design around a stable session ID and context that lives across invocations, not stateless handlers that rehydrate everything each turn.
  2. Pick the compute shape from the traffic shape. Spiky and latency sensitive goes on microVMs; continuous, GPU bound, or multi agent goes on instances. The APIs are the same, so switching costs are low.
  3. Separate ephemeral session state from durable state. Use the session filesystem for working data, AgentCore Memory and EBS for anything that must outlive it.
  4. Do not confuse long sessions with durable workflows. For human approvals, money movement, or anything needing transactional recovery, put a Step Functions orchestrator in charge and make the agent a step.
  5. Watch idle cost on instances. No scale to zero means you stop sessions deliberately.

The deeper story is that agents are a genuinely new compute primitive, and infrastructure is catching up to that fact. For a decade we bent every workload to fit the web request. AgentCore is a bet that it is time to bend the compute to fit the workload instead. If your agents still look like HTTP handlers in disguise, that is probably the first thing to rethink.

Years ago I sat with a customer who needed to orchestrate multiple agents inside their software products — agents that coordinate, hand work to each other, and stay alive across a workflow, not one model answering one call. We wrote a PRFAQ together, took it into an EBC, and a...