The Principal's Filter: Sorting Agent Hype from What Ships

Last month a customer walked me through a slide deck with eleven “agents” on it. By the end of the hour we had found one real agent, three workflows, two chatbots, and five ideas that nobody had scoped yet. Nobody in that room was being dishonest. The word “agent” has simply stretched to cover almost anything with a model behind it.

That is the real job of a principal right now. Customers do not need another person cheering for agents. They need someone who can sort what is viable from what is not, then help them build the right thing, which is sometimes an agent and often is not.

Why This Matters Now

The gap between buzz and production is wide. A Gartner 2026 CIO survey found that only 17% of organizations have deployed agents, while more than 60% expect to within two years. The Sinequa State of Enterprise Agentic AI 2026 report is even more sobering: only 24% have deployed a true agent, and only 10% have true agentic capabilities. Gartner also forecasts that more than 40% of agentic AI projects will be canceled by the end of 2027 because of rising cost, unclear value, and weak risk controls.

Notice what is missing from that list: model quality. In my experience, most agent projects do not fail because the model is weak. They fail because of how they are governed and run. The value was never clear, the costs crept up, the risk controls were thin, and nobody was assigned to manage the agent once it went live.

So I use a filter. Five questions, asked early, before anyone writes a line of orchestration code.

Filter 1: Is It Really an Agent?

The industry now has a name for the relabeling problem: agent washing. Chatbots, assistants, and RPA tools are being renamed as agents, and buyers are paying agent prices for them.

My test is simple. A real agent does three things:

  • Plans. It breaks a goal into steps it was not explicitly given.
  • Uses tools. It calls APIs, queries data, or takes actions in other systems.
  • Acts on its own. It decides the next step based on what it just observed, without a human or a fixed script choosing for it.

If the steps are known in advance, you have a workflow. Build it as a workflow. A state machine with a model call in one or two steps is cheaper, easier to test, and easier to explain to an auditor. I tell customers this is not a downgrade. It is the correct design.

Fixed steps, known branches   -> Workflow (Step Functions, plus a model call)
Answer questions, no actions  -> Assistant / RAG
Repetitive UI clicks          -> RPA
Open ended goal, tool choice  -> Agent

Filter 2: What Happens When It Fails?

Every agent demo works. That is what demos are for. The question is what happens on day thirty when the input looks different. Agents that shine in demos break on new input formats, loop on vague instructions, or quietly make up data. The last one is the dangerous one, because a confident wrong answer does not trigger an alarm.

I ask the team to describe three failure paths out loud:

  1. Catching it. How do you know the agent failed? Schema validation on outputs, checks against a system of record, and confidence thresholds that route to a human.
  2. Containing it. What is the blast radius? Scoped credentials, read only access by default, step and spend limits per task, and an approval gate before any irreversible action.
  3. Recovering from it. Can you replay the trace, see which tool call went wrong, and roll back what the agent did?

If the answer to any of these is “the model is pretty good, so it should be fine,” the project is not ready.

Filter 3: Can You Measure and Test It?

You cannot improve what you cannot measure, and you cannot ship what you cannot test. Before production, I want to see an evaluation suite built on the customer’s real edge cases, not the vendor’s happy path examples.

A useful starting point is fifty to two hundred real tasks pulled from tickets, logs, or past work, each with a known good outcome. Run the agent against them on every prompt change, model change, and tool change. Something as plain as this works:

results = []
for case in eval_cases:
    outcome = agent.run(case["input"], max_steps=15)
    results.append({
        "id": case["id"],
        "passed": grade(outcome, case["expected"]),
        "steps": outcome.step_count,
        "cost_usd": outcome.cost_usd,
    })

pass_rate = sum(r["passed"] for r in results) / len(results)
print(f"pass rate: {pass_rate:.1%}")

Then carry the same idea into production. Sample live traces, grade them, and track pass rate, step count, and escalation rate over time. Offline evaluation tells you whether to ship. Production monitoring tells you when to stop.

Filter 4: What Does Each Completed Task Cost?

Token pricing is the wrong unit. Customers get excited about a fraction of a cent per call, then discover that one finished task took nine calls, two retries, a reasoning loop, and fifteen minutes of a person checking the result.

The number I ask for is cost per completed task:

cost per completed task =
  (model calls + tool calls + retries + loops
   + infrastructure + human review time)
  / tasks completed correctly

The denominator matters as much as the numerator. If the agent completes 70% of tasks correctly, the other 30% still cost money and still need a person to finish them. Compare the result to what the work costs today. If the agent is not clearly cheaper, faster, or better at the same quality bar, the value case is not there yet. This is exactly the “rising cost, unclear value” pattern behind those cancellation forecasts.

Filter 5: Who Manages It Once It Is Live?

This is the question that kills the most projects, and the one teams skip most often. An agent is not a feature you ship and forget. It behaves more like a new team member who needs a manager.

Someone has to:

  • Define the work. What tasks are in scope, and which are not.
  • Set the quality bar. What counts as done, and what error rate is acceptable.
  • Step in when things go wrong. Review escalations, pause the agent, and fix the prompt, tools, or data.
  • Own the evaluation suite. Add new edge cases as production reveals them.

I ask for a name, not a team. If no single person owns that job, with time actually allocated to it, the project is not viable, regardless of how good the demo looked.

Practical Takeaways

  1. Classify before you build. Label every proposed “agent” as a workflow, assistant, RPA, or agent. Most will not be agents.
  2. Write the failure paths first. Catch, contain, and recover, documented before the first sprint.
  3. Build the evaluation suite from real data. Use the customer’s edge cases and rerun it on every change.
  4. Price the finished task, not the token. Include retries, loops, and human review.
  5. Name the owner. No operational owner, no production launch.

Saying “Not Yet” Is the Job

Principals earn trust by being the person in the room willing to say “not yet,” or “a simpler build will work better.” Customers remember who saved them from an expensive cancellation far longer than who sold them the shiniest demo.

And when the filter says an agent really is the answer, the right primitives exist. As I wrote in Agents Aren’t Web Requests, agents need runtimes built for long running, stateful, isolated work. Get the governance right first, and the infrastructure part becomes the easy part.

What questions are in your filter? I would like to hear what you ask customers before you let an agent near production.

Last month a customer walked me through a slide deck with eleven “agents” on it. By the end of the hour we had found one real agent, three workflows, two chatbots, and five ideas that nobody had scoped yet. Nobody in that room was being dishonest. The word “agent” has...

Agents Aren't Web Requests — Why AWS Rebuilt the AgentCore Runtime

Years ago I sat with a customer who needed to orchestrate multiple agents inside their software products — agents that coordinate, hand work to each other, and stay alive across a workflow, not one model answering one call. We wrote a PRFAQ together, took it into an EBC, and a VP agreed to build it. The document was right. It was also complex — it described, in detail, nearly every capability you can find in AgentCore today. The service team then rewrote it, more than once, not to correct it but to break a correct and complex vision into something they could actually execute and build. The primitives got names and edges they did not have in our draft. But the shape of the thing never changed, because we had the mental model right from the start: these are not web requests. They are sessions.

That is the part I am proud of, and it is the only part that matters here: the vision was correct and complete, and the one idea holding all of it together was that an agent is a session, not a request.

Here is the mental model that quietly breaks the moment you ship an agent to production: “an agent is just a web request that happens to call an LLM.” You wire up an HTTP handler, it fires off a prompt, you stream tokens back, the function returns, the compute is reclaimed. Clean. Stateless. Familiar.

It is also wrong in ways that will cost you. An agent is not a request. It is a process — one that reasons over many turns, accumulates intermediate results, calls tools that mutate state, occasionally reaches for a GPU, and sometimes hands work to other agents. The shape of that workload has almost nothing in common with the shape of an HTTP handler. Amazon’s answer, Amazon Bedrock AgentCore Runtime, is best understood not as “serverless with a bigger timeout” but as a deliberate rethinking of what compute for agents should look like.

What breaks when you run an agent like an HTTP handler

Serverless web infrastructure makes three assumptions that are load bearing for request response traffic and actively hostile to agents.

It assumes statelessness. An HTTP handler is supposed to carry nothing between invocations; that is what lets the platform scale it horizontally without thinking. But an agent’s whole value is the state it carries — conversation history, a scratchpad of intermediate reasoning, the output of a tool it called two steps ago, a half built file on disk. If every invocation lands on a fresh execution environment with an empty filesystem, you are forced to serialize and rehydrate the entire context on every single turn. That is slow, lossy, and fragile.

It assumes short duration. Request response compute is tuned for work that completes in milliseconds to seconds, with hard timeouts measured in minutes. Agents routinely run for minutes to hours, and some genuinely need to run continuously for days — a transformation job grinding through a corpus, an automation loop that pauses and resumes, a research agent that keeps working while a human is asleep.

It assumes weak isolation is fine. When requests are stateless and ephemeral, co tenancy is mostly an efficiency detail. But agents run privileged, nondeterministic code on a user’s behalf — they get shell access, read and write files, and invoke tools with real credentials. The isolation boundary stops being a performance concern and becomes a security boundary. You cannot have user A’s agent able to observe anything left behind by user B’s.

How AgentCore answers it: isolated sessions, not shared handlers

AgentCore Runtime’s primitive is the session, not the request. Each user session runs in its own dedicated microVM with isolated compute, memory, and filesystem, plus shell access. One user’s agent cannot reach another user’s data, and when the session ends the entire microVM is torn down and memory is sanitized — no cross session residue. That is a deterministic isolation boundary wrapped around a deliberately nondeterministic workload, which is exactly the property enterprises need.

Inside that boundary, the model does the thing HTTP handlers refuse to do: it lets you safely reuse context across invocations. You generate a session ID, pass the same ID on every related call, and each invocation builds on the environment the previous one left behind — same filesystem, same in memory state, same tool context. Multi turn conversations and multi step workflows stop being a serialization problem and become what they always should have been: successive calls into a living environment.

The honest part of the design is that this state is ephemeral by default. In memory and on disk data lives for the session’s lifetime and no longer. For anything that must outlive the session — learned preferences, durable conversation history, workspace files — you pair the runtime with AgentCore Memory for long term recall and Amazon EBS for persistent volumes. The runtime handles scaling, session management, isolation, patching, and observability so you spend your time on agent logic, not plumbing.

Two compute shapes, one set of APIs

The sharpest architectural decision is that AgentCore exposes two compute shapes under the same runtime APIs, so you match the shape of the compute to the shape of the traffic.

Serverless microVMs are the default: fast cold starts, scale to zero, sessions up to 8 hours. This is the right home for request response, spiky, or latency sensitive agents — the chat assistant, the API driven tool, the thing that needs to start instantly and cost nothing while idle.

Runtime instances, generally available since August 6 2026, are AWS managed EC2 running in your own account, defined by a reusable capacity provider (operating system, allowed instance types, networking, storage, IAM roles). Sessions here run up to 14 days, support GPU instance families with drivers provisioned for you, and can stop and restart to save cost — hibernate Monday night, resume Wednesday morning with volumes reattached and data intact. This is the home for long running, continuous, GPU bound, or multi agent work.

Because the instances live in your account, your data and your account controls stay put, and you can apply existing Savings Plans, Reserved Instances, and ODCRs. AgentCore still owns the lifecycle — provisioning, patching, scaling, teardown — so you get EC2 economics without EC2 babysitting.

Multiple agents on one host

Runtime instances change the unit of collaboration. In the microVM model, one runtime hosts one agent. On an instance, a single session can host many agents that share a filesystem. AWS’s own demo makes the point: a writer agent generates code into a shared session directory, and a reviewer agent reads that same file and critiques it — no message passing, no inter agent API calls, no data transfer. They collaborate through the filesystem.

That composes cleanly with the two shapes. A lightweight orchestrator on a microVM can take fast, spiky, API driven traffic and dispatch heavier work to specialized worker agents running on instances — code compilation, security scanning, GUI automation — that need persistent state and direct OS access. And none of this locks you into a framework: bring CrewAI, LangGraph, LlamaIndex, or Strands, and any model. Packaging is minimal — an @app.entrypoint decorator plus a zip or a container image.

The caveat worth saying out loud

Here is where I will push back on the easy narrative. A 14 day session is not a durable workflow. If your process waits days for a human approval, or moves money, or must survive infrastructure failure with transactional guarantees, a single long lived session is a fragile place to park that state. Runtime instances raise the ceiling for continuous work; they do not turn one invocation into a saga. For orchestration that must be durable and recoverable, reach for a real orchestrator — AWS Step Functions over a transactional store — and let the agent be a step inside it, not the system of record.

Two more things to keep honest. Instances do not scale to zero; an idle instance still bills until you stop the session. And the whole point of the two shapes is that neither is universally correct. Match the compute shape to the traffic shape or you will overpay for idle capacity or starve a long job of the persistence it needs.

Takeaways

  1. Model agents as sessions, not requests. Design around a stable session ID and context that lives across invocations, not stateless handlers that rehydrate everything each turn.
  2. Pick the compute shape from the traffic shape. Spiky and latency sensitive goes on microVMs; continuous, GPU bound, or multi agent goes on instances. The APIs are the same, so switching costs are low.
  3. Separate ephemeral session state from durable state. Use the session filesystem for working data, AgentCore Memory and EBS for anything that must outlive it.
  4. Do not confuse long sessions with durable workflows. For human approvals, money movement, or anything needing transactional recovery, put a Step Functions orchestrator in charge and make the agent a step.
  5. Watch idle cost on instances. No scale to zero means you stop sessions deliberately.

The deeper story is that agents are a genuinely new compute primitive, and infrastructure is catching up to that fact. For a decade we bent every workload to fit the web request. AgentCore is a bet that it is time to bend the compute to fit the workload instead. If your agents still look like HTTP handlers in disguise, that is probably the first thing to rethink.

Years ago I sat with a customer who needed to orchestrate multiple agents inside their software products — agents that coordinate, hand work to each other, and stay alive across a workflow, not one model answering one call. We wrote a PRFAQ together, took it into an EBC, and a...

When Building Gets Cheap, Deciding Gets Expensive - The New Engineering Bottleneck

When I led development teams, I always believed that the ability to build was our edge. Give the team room and flexibility and we could make almost anything ourselves. But building was never the real question. Every serious feature came down to build versus buy, and the deciding factor for me was always the same: was this core business logic that actually differentiated us? If it was, we built it and owned it. If it was not, buying was almost always the smarter call. What has changed is not that logic. It is that the cost of building has fallen so far that the temptation to build everything is stronger than ever, even when it is not the right decision.

Last month I watched a mid level engineer stand up a working service in about forty minutes. Authentication, a REST layer, input validation, a test suite, and a passable README. Two years ago that was a sprint. What struck me was not the speed. It was the silence in the room afterward, because the hard conversation had not even started. Should this service exist? Should it own that data? Was REST the right call, or did we just build the thing the assistant was fastest at generating?

That gap between “we can build it” and “we should build it, this way, now” is where engineering work is quietly relocating. For most of our careers, implementation was the expensive part. That assumption no longer holds, and a lot of team design is built on top of it.

The old bottleneck is collapsing

The cost of turning an idea into running code has been falling for a long time. High level languages, package managers, cloud primitives, and copy paste from the wider internet all chipped away at it. AI coding assistants and agentic workflows have turned that slow decline into a cliff. A prototype that used to take a week takes an afternoon. A one off migration script that used to take a day takes ten minutes.

When you cut the price of something dramatically, you get more of it. Teams are now generating more candidate implementations, more branches, more prototypes, and more “what if we just tried it this way” experiments than ever before. That is genuinely good. Cheap building means cheap exploration.

But cost does not disappear. It moves. And it has moved to the one part of the pipeline that no assistant has automated away: deciding.

When building gets cheap, deciding gets expensive

Consider what actually happens now when a team picks up a feature. The assistant can produce three plausible designs before lunch. Now someone has to choose. Which data model do we commit to? Do we accept the coupling this introduces? Is this a change we can walk back next quarter, or are we about to bake it into a public API that a dozen downstream teams will build on?

Those questions were always there. What changed is their share of the total effort. When building was 80 percent of the work, deciding was a rounding error you could absorb in a hallway conversation. When building drops to 20 percent, deciding becomes the main event, and most teams have no process for it beyond that same hallway conversation, now hopelessly overloaded.

This is the shift in one sentence: the scarce resource is no longer the ability to produce a working artifact. It is the judgment to know which artifact is worth keeping and what it commits you to.

Why decisions are now the constraint

Three forces make decisions the binding constraint rather than just a cost that grew.

More options per unit time. Every generated alternative is a decision you now owe. Ten prototypes is not ten times the progress. It is ten times the deciding, and deciding does not parallelize the way generation does. A single senior engineer or architect becomes the queue that every branch waits in.

Decision debt compounds quietly. We talk endlessly about technical debt. Decision debt is worse because it is invisible until it detonates. It is the accumulation of choices made implicitly, by default, or by whoever’s code got merged first. When building was slow, the pace of code creation throttled how fast you could accrue this debt. Remove the throttle and you can bury a team in unexamined commitments in a single quarter.

Alignment becomes the real work in progress. Watch where things actually stall today. Rarely at the keyboard. Usually in the review queue, the design doc waiting for comments, the Slack thread where three staff engineers disagree about an approach and nobody owns the tiebreak. That is decision latency, and it is now the dominant term in your cycle time.

What good looks like

The teams pulling ahead are the ones treating decision making as a first class engineering discipline, not an afterthought. A few patterns are doing real work.

Classify the door before you walk through it. The most useful frame I know separates reversible decisions from irreversible ones. Cheap to reverse? Do not convene a committee. Let an engineer pick, ship, and learn. Expensive or impossible to reverse, like a public API contract, a data retention model, or a security boundary? That earns real deliberation. The failure mode I see most is teams applying heavyweight process to reversible choices and no process at all to the irreversible ones. When building is cheap, most decisions become more reversible, which means you should be pushing far more of them down and out.

Write the decision down, lightly. A short architecture decision record captures the context, the options, the choice, and the tradeoff accepted. Two paragraphs beats a forty page design doc nobody reads. The point is not ceremony. It is that six months from now, when someone asks “why is it like this,” the answer exists and does not require archaeology. In a world where the code was half generated, the reasoning behind it is the artifact worth preserving.

# ADR 042: Event sourcing for the ledger service

Status: Accepted
Context: We need an auditable history of balance changes.
Decision: Append only event log as the source of truth.
Tradeoff accepted: Higher read complexity now, in exchange for
  a complete audit trail we cannot reconstruct later.
Reversibility: One way door. Revisit only with a migration plan.

Assign the owner, not the committee. Every consequential decision needs a name attached, someone accountable for making the call and living with it. Consensus feels safe and is usually just diffused ownership. Fast teams name a decider, gather input on a clock, and move.

Attack decision latency directly. Measure how long a design sits in review. Put a timebox on it. Make “we decided” a visible event, not something that emerges by attrition when the loudest person tires out.

Implications for technical leaders and topologies

If deciding is the new bottleneck, the job of a technical leader changes. The highest leverage senior person on a team used to be the one who wrote the trickiest code. Increasingly it is the one who makes the tricky calls fast and well, and who builds the machinery for the rest of the team to do the same.

That means leaders shift from code reviewers to decision enablers. Reviewing every line does not scale when the line count explodes. What scales is establishing the guardrails, the defaults, and the patterns that let engineers decide safely on their own. Think guardrails over gatekeeping. A paved road that makes the safe choice the easy choice removes a thousand small decisions from the queue before they ever form.

It also means we invest deliberately in judgment and taste, the things assistants do not have. Knowing that the elegant abstraction is premature. Sensing that a dependency will hurt in a year. Feeling when a design is too clever. That intuition is now the differentiated skill, and it is developed by giving engineers real decisions to own early, then reviewing the outcomes together.

The edge is decision velocity and decision quality

The comfortable metric for a decade was output. Story points, PRs merged, features shipped. When output was expensive, output was a fine proxy for progress. It is not anymore. A team that generates twice the code but takes three weeks to decide what is worth keeping is slower than a team that generates half as much and decides in a day.

The competitive edge is moving to decision velocity multiplied by decision quality. Deciding fast, deciding well, and knowing which decisions deserve which level of care. Building is nearly free now. What you choose to build, and who gets to choose, is the whole game.

Further reading worth your time: the reversible and irreversible decisions framing, Michael Nygard’s original ADR write up, and the Team Topologies material on cognitive load and team design.

When I led development teams, I always believed that the ability to build was our edge. Give the team room and flexibility and we could make almost anything ourselves. But building was never the real question. Every serious feature came down to build versus buy, and the deciding factor for...

Stop Reaching for the Biggest Model

Stop Reaching for the Biggest Model

There is a reflex that shows up in nearly every generative AI design review I sit in. Someone needs to classify support tickets, or extract fields from invoices, or summarize a document, and the default choice is whatever frontier model is topping the leaderboard this quarter — Claude Opus, GPT-5.x, Gemini. It is the obvious pick. It is the safe pick. And it is frequently the wrong pick.

The problem is not that frontier models are bad. They are extraordinary. The problem is that you are about to multiply that per call premium by millions of calls, and nobody in the room has asked whether the job actually requires a frontier model at all. The right question is almost never “what is the best model?” It is “what is the cheapest architecture that clears my quality threshold reliably?”

Those are very different questions, and they lead to very different bills.

The economics nobody runs the numbers on

Let me make this concrete with a workload I see constantly: summarization at scale. Suppose you run 10,000 summaries a day. Each request is roughly 1,000 input tokens and 300 output tokens — a realistic shape for document or thread summarization.

Run that on a small, task specific model in the 2B to 14B parameter range and you are looking at something like $7.20 a day. Run the identical workload on a mid tier frontier option like GPT-4.1 Mini and you land closer to $58 a day. Same inputs, same outputs, and — this is the part that matters — in the benchmark those numbers come from, human evaluators could not detect a quality difference in the summaries themselves.

That gap is about $18,500 a year on a single workload. Not your whole platform. One job. Most organizations have a dozen of these running quietly in production, each one silently overpaying because the model choice was made once, early, by reflex, and never revisited.

The AWS Machine Learning blog made this point sharply in Beyond the price per token: the sticker price on a token is not the number that matters. What matters is the cost of a completed, acceptable task. And when the task is narrow and measurable, a smaller model routinely wins that comparison outright — lower cost, lower latency, and quality that is indistinguishable from the expensive option.

The honest counterexample

Now, if I stopped here I would be selling you a lie by omission. Small does not mean smart. Small means small.

Consider reasoning. On GSM8K — a grade school math word problem benchmark that demands multi step reasoning — a 0.5B parameter model scored around 37.7%. Llama-3.1-70B scored 92.0% on the same benchmark. That is not a rounding error. That is the difference between a system you can ship and a system that is wrong on nearly two thirds of its answers.

So this is not “small always wins.” This is task matching. The summarization job and the math reasoning job look superficially similar — both are “send text, get text” — but they place completely different demands on the model. Summarization is largely a compression and rephrasing task that small models handle beautifully. Multi step arithmetic reasoning is exactly the capability that scales with parameter count, and starving it produces garbage.

The discipline is knowing which is which. High volume, narrow, well defined work — classification, extraction, routing, summarization, format conversion — is small model territory. Open ended reasoning, complex code generation, nuanced judgment across long context — that is where you pay for the frontier, and you should. Both Refonte and Forbes have documented small language models beating frontier systems on cost, speed, and accuracy — but always on the tasks that fit them.

The discipline: measure completed task economics

If you take one thing from this post, take this: stop comparing API price cards. Price per token is an input, not an outcome. What you actually pay for is a completed task that meets your bar, and that number includes things the pricing page never shows you:

  • Accuracy at your threshold. Not headline benchmark scores — accuracy on your data, at your acceptable quality line.
  • Retries. A cheaper model that fails and gets rerun twice is not cheaper.
  • Human review rate. If a smaller model pushes 15% of outputs to a human reviewer and a larger one pushes 3%, the reviewer’s time may dwarf any token savings.
  • Escalation cost. What does a wrong answer cost downstream when it slips through?

When you sum those up, the winning model is often not the one with the lowest token price or the highest benchmark score. It is the one that clears your threshold with the least total cost per acceptable output.

The AWS Well-Architected Generative AI Lens codifies this as GENCOST01-BP01: Right-size model selection. The guidance is refreshingly blunt — start with the smallest model that could plausibly work, and scale up only when your evaluations show you need to. This is the inverse of the reflex. Instead of starting at the top and hoping you can justify the cost, you start at the bottom and earn your way up with evidence.

A pattern that makes this operational is complexity based routing. Rather than sending every request to one model, you inspect the request and route it: simple ones to a small fast model, hard ones to the frontier. On AWS you can do this with Amazon Bedrock Intelligent Prompt Routing, which dynamically dispatches each prompt to the most cost effective model that can handle it.

incoming request
      |
      v
  classify complexity  ──> simple / narrow  ──> small task specific model (2B–14B)
      |                                              |
      |                                              v
      |                                        meets threshold? ── yes ──> done
      |                                              | no
      v                                              v
  complex / open-ended ──────────────────> frontier model

And critically: right sizing is not a one time decision. New models ship every few weeks, and the price to performance frontier moves under you constantly. Treat model selection as a continuously revisited parameter, not a setting you configure once and forget.

Practical takeaways

  1. Ask the right question. Not “what is the best model?” but “what is the cheapest architecture that clears my quality threshold reliably?”
  2. Measure completed task economics, not price per token. Fold in retries, human review rate, and escalation cost.
  3. Start small and scale up. Follow GENCOST01-BP01: begin with the smallest plausible model and only move up when evaluations demand it.
  4. Route by complexity. Use Bedrock Intelligent Prompt Routing so simple requests get small model pricing and hard ones still get frontier quality.
  5. Re-evaluate continuously. The efficient frontier moves. Revisit your model choices as new options ship.

Paying for general purpose capability you never use is not a convenience. It is an architectural choice — and for high volume, narrow, measurable work, it is usually the wrong default…

Stop Reaching for the Biggest Model

There is a reflex that shows up in nearly every generative AI design review I sit in. Someone needs to classify support tickets, or extract fields from invoices, or summarize a document, and the default choice is whatever frontier model is topping the...

Can AI Actually Run Your On-Call? The AI SRE Reality Check

It is 3:14 AM. The pager fires. Latency on the checkout service is climbing, error budgets are burning, and a dashboard somewhere has turned an angry shade of red. The pitch from a dozen vendors this year is seductive: what if no human had to wake up at all? What if an AI SRE triaged the alert, correlated the signals, ran the remediation, and closed the incident before you finished your first REM cycle?

It is a great demo. It is also, in most production environments, not yet true. The question worth asking is not “can AI run my on-call” as a binary, but “what parts of on-call can I safely hand to AI today, and what parts still demand a human on the hook?” The answer is more encouraging than the skeptics claim and more sobering than the marketing decks suggest.

The Pitch Versus the Reality

The strongest version of the “AI runs your on-call” claim imagines an autonomous agent that holds the pager, owns the rotation, and makes remediation decisions without a human in the loop. That framing conflates two very different things: assisting the responder and replacing the responder.

The gap between them is accountability. When an incident causes data loss or a multi hour outage, someone answers for it — in the postmortem, to the customer, and sometimes to a regulator. An AI agent cannot be that someone. It has no career at stake, no context outside the telemetry it was fed, and no standing to accept the risk of a destructive remediation. That single constraint reshapes everything about how AI fits into on-call. The technology can do an enormous amount of the work. It cannot hold the responsibility.

So the useful question becomes architectural: where does AI genuinely reduce toil and time to resolution, and where does inserting it create a false sense of safety?

What AI SRE Genuinely Handles Well Today

Strip away the autonomy fantasy and a lot of real value remains. The tasks AI handles well share a common shape: they are high volume, pattern rich, and low blast radius. Getting them wrong wastes minutes, not customers.

Alert triage and noise reduction. The single biggest source of on-call misery is alert fatigue. A model that has ingested months of alert history can cluster related alerts, suppress known flapping signals, and rank what actually deserves a human’s attention. This is a classification problem with abundant labeled data, and it is exactly where machine learning shines. Cutting a 200 alert storm down to the three that matter is not a party trick — it is the difference between a responder who thinks clearly and one who drowns.

Signal correlation across telemetry. Modern systems emit metrics, logs, traces, and events across dozens of services. A human under pressure can hold maybe a handful of these in working memory. An AI assistant can correlate a latency spike with a recent deploy, a spike in a downstream dependency, and an anomalous log pattern, then surface the join in seconds. It is doing the tedious cross referencing that a senior engineer would do, faster, and without fatigue.

Drafting the incident timeline. Writing up what happened and when is real work that usually gets deferred until the postmortem, by which point the details are fuzzy. An AI that watches the incident channel and the telemetry can draft a coherent timeline in real time — deploy at 03:02, error rate inflection at 03:07, rollback initiated at 03:19. The responder edits rather than reconstructs.

Suggesting runbook steps. When the failure mode is known and documented, retrieving the right runbook and proposing the next step is a retrieval and ranking task that AI does well. “This looks like the connection pool exhaustion we saw in March; the runbook says bump the pool size and recycle the workers” is a genuinely useful suggestion, provided a human confirms it.

Notice the pattern. In every one of these, the AI compresses information and proposes action. A human still decides. That division is not a limitation to engineer away — it is the design.

Where It Still Needs a Human

The tasks AI struggles with share the opposite shape: they are rare, ambiguous, and high stakes. These are precisely the incidents that define your reliability, and precisely where the confident answer is the dangerous one.

Novel failure modes. AI is fundamentally an interpolation engine. It is superb at “this looks like something I have seen before” and unreliable at “this has never happened.” The gnarly outages — a cascading failure from an unexpected interaction between two services, a corruption bug that only manifests under a specific load pattern — are novel by definition. There is no training example, and the model will often pattern match to the nearest familiar incident, which is the wrong one. A human recognizes “this is weird” in a way current systems do not.

Ambiguous blast radius. Deciding how far a problem reaches, and how far a fix reaches, requires a mental model of the whole system plus business context the telemetry does not contain. Is this degrading a background job or the payment path? Is the affected customer a free tier user or the account that renews next week? AI sees the graph of services. It rarely sees the map of what actually matters.

High stakes remediation. There is a category of action you cannot take back: failing over a database, dropping traffic, deleting state, scaling a fleet to zero. The cost of a wrong autonomous decision here is not minutes, it is a second, worse incident on top of the first. Every mature on-call practice puts a human decision gate in front of destructive actions, and AI does not change that calculus. Let the AI propose the failover. Let a human approve it.

Accountability and ownership. Even if the model were right every time, someone still has to own the outcome. Ownership is what motivates the careful judgment, the pre incident hardening, and the honest postmortem. Diffuse that onto a tool and you erode the culture that made the system reliable in the first place.

A Practical Adoption Model

The pragmatic path is not “replace the rotation” and it is not “ignore the technology.” It is to treat AI as a first responder assistant that makes your existing humans faster and calmer. A staged model works well:

  1. Deploy AI on triage and correlation first. These are low risk, high value, and build trust. Measure the reduction in alert volume and in time to first meaningful signal.
  2. Let AI draft, humans decide. Timelines, runbook suggestions, and remediation proposals flow through the AI, but a human confirms every action that touches production state.
  3. Gate destructive actions behind explicit human approval, always. No autonomous failovers, deletes, or fleet wide changes. This is a hard line, not a phase to graduate out of.
  4. Keep the human on the pager. The AI reduces how often the pager fires and how hard each page is to resolve. It does not hold the pager. Accountability stays with a named person.
  5. Feed every incident back in. The AI gets better as your incident history grows. Treat postmortems as training data and the assistant compounds in value over time.

The Honest Take

Can AI run your on-call? No — not if “run” means owning the pager and making autonomous high stakes calls. Can AI make your on-call dramatically less painful, faster, and more consistent? Absolutely, and if you are not already piloting it on triage and correlation, you are leaving real toil reduction on the table.

The winning teams are not the ones chasing a fully autonomous SRE that does not exist yet. They are the ones deploying AI as a force multiplier: a tireless assistant that handles the volume so the human can handle the judgment. Keep a person on the hook, put a human gate in front of anything you cannot undo, and let the machine do the rest. That is not a compromise. Right now, it is the state of the art.

It is 3:14 AM. The pager fires. Latency on the checkout service is climbing, error budgets are burning, and a dashboard somewhere has turned an angry shade of red. The pitch from a dozen vendors this year is seductive: what if no human had to wake up at all? What...

Agentic RAG Is Just Retrieval Growing a Spine

Ask a naive RAG system a question like “Which of our regions missed their SLA last quarter, and what caused it?” and watch it faceplant. It embeds the whole sentence, pulls the top eight chunks that look vaguely similar, stuffs them into a prompt, and hopes the model can reason its way out. But the answer lives in two different documents — one listing SLA breaches, another explaining root causes — and neither is close enough to the query vector to rank in the top eight. The model gets half the context, confidently fabricates the rest, and you ship a wrong answer with a citation attached.

This is the structural limit of one-shot retrieval. Agentic RAG fixes it not by tuning the embedding model or cranking up k, but by giving retrieval something it never had: the ability to decide. That decision-making capacity is the spine.

The limits of one-shot RAG

Classic RAG is a straight pipeline. Embed the query, run a nearest-neighbor search, take the top k chunks, concatenate, generate. It works beautifully for questions that map cleanly to a single passage — “What is the default timeout for a Lambda function?” — and falls apart everywhere else.

The failure modes are predictable once you see the pattern:

  • Compositional questions. Answers that require joining facts across documents. Top-k similarity has no notion of “I need one fact from here and another from there.”
  • Vague or underspecified queries. “How do we handle failures?” retrieves everything and resolves nothing. The query needs sharpening before it can retrieve anything useful.
  • Multi-hop reasoning. “Who approved the change that caused the outage?” requires finding the outage, then the change, then the approver — three sequential lookups, each depending on the last.
  • No self-awareness. One-shot RAG cannot tell whether what it retrieved is any good. It retrieves once and commits, whether the chunks are relevant or garbage.

You can paper over some of this with hybrid search, reranking, and bigger context windows. Those help. But they are all still one shot. The system fetches, then generates, and never asks itself whether it fetched the right thing.

What “agentic” actually adds

Agentic RAG wraps retrieval in a control loop. Instead of a fixed fetch-then-generate sequence, the model reasons about the task and treats retrieval as a tool it can invoke — repeatedly, conditionally, and with different arguments each time.

The loop looks roughly like this:

while not satisfied and steps < budget:
    plan  = model.decide_next_action(question, evidence_so_far)
    if plan.action == "retrieve":
        results  = knowledge_base.retrieve(plan.query)
        evidence += model.filter_relevant(results)
    elif plan.action == "answer":
        break
answer = model.generate(question, evidence)

Four decisions live inside that loop that one-shot RAG never makes: whether to retrieve at all (some questions do not need it), what to retrieve (the model writes the search query, not the user), how many times to retrieve, and when the accumulated evidence is good enough to stop. Retrieval stops being a passive lookup and becomes an active, self-directed component. That is what “growing a spine” means — the system now holds itself upright and makes calls instead of flopping through a fixed pipeline.

Core patterns

Three patterns do most of the heavy lifting in production agentic RAG.

Query reformulation. The user’s phrasing is rarely the best search query. An agentic system rewrites it — expanding acronyms, splitting a compound question into parts, or generating several candidate queries and merging the results. When a retrieval comes back weak, it reformulates and tries again rather than committing to bad context.

Multi-hop retrieval. For questions that require chaining, the agent retrieves, reads, and uses what it learned to form the next query. Answering “what caused the SLA miss” becomes: retrieve the breach record, extract the incident ID, retrieve the incident report, extract the root cause. Each hop is a fresh, better-targeted search.

Self-critique and grounding checks. This is the idea behind research like Self-RAG and Corrective RAG (CRAG). The model grades its own retrievals for relevance and grades its draft answer for grounding — is every claim actually supported by a retrieved passage? If a claim is unsupported, it retrieves again or drops the claim. This is the single biggest lever against hallucination, because the system refuses to assert what it cannot cite.

Production realities

A control loop that can retrieve N times is a control loop that can retrieve N times when N gets ugly. The engineering discipline matters more than the pattern.

Latency compounds. Each hop is a full round trip: an LLM call to plan, a vector search, another LLM call to evaluate. A three-hop answer can mean six or seven sequential model invocations. Where hops are independent, fan them out in parallel. Where they are sequential, use a smaller, faster model for the planning and grading steps and reserve the large model for the final synthesis.

Token cost accumulates. Every hop re-sends accumulated evidence into the context window. Naively, a five-hop conversation can burn several times the tokens of one-shot RAG. Summarize intermediate evidence instead of carrying raw chunks forward, and cap how much context each hop contributes.

Guard the loop. Never ship an unbounded while. Enforce a hard ceiling on iterations, a token budget for the whole request, and a wall clock timeout. Track a confidence or novelty signal — if two consecutive hops add nothing new, stop. A runaway agent that retrieves forty times is worse than a wrong answer, because it is a wrong answer that also costs forty dollars.

Evaluate retrieval separately from generation. Measure retrieval quality (recall, precision, whether the right passages showed up) independently from answer quality (faithfulness, correctness). A good final answer built on lucky guesses is a landmine. Build a labeled question set and track both metrics as you tune. Frameworks like Ragas exist precisely for this split.

Building it on AWS

You do not have to hand-roll the whole loop. AWS gives you the pieces at two levels of abstraction.

The managed retrieval layer. Amazon Bedrock Knowledge Bases handles ingestion, chunking, embedding, and vector storage, and exposes a Retrieve API for raw passages and a RetrieveAndGenerate API for one-shot answers. In an agentic design you lean on Retrieve as the tool your loop calls — the agent owns the reasoning, the Knowledge Base owns the fetch.

The managed orchestration layer. Agents for Amazon Bedrock runs the reason-act loop for you. You attach a Knowledge Base and define action groups, and the agent plans, decides when to query, invokes tools, and iterates — emitting a reasoning trace so you can see why it retrieved what it did. That trace is not a nicety; it is how you debug and audit a multi-hop answer.

The build-versus-buy line is straightforward. If your loop is standard retrieve-reason-retrieve, let Agents orchestrate it and save yourself the state machine. When you need custom stopping logic, parallel fan-out, or tight control over every model call, drop to the Retrieve API and orchestrate yourself — with Step Functions for durable multi-step flows or a framework like LangGraph running on Lambda or ECS for tighter loops.

Closing take

Agentic RAG is not a wholesale replacement for classic RAG — it is classic RAG that finally learned to think about what it is doing. And it is not free. Every hop costs latency, tokens, and complexity, so reach for it when your questions genuinely demand it: compositional queries, multi-hop reasoning, high-stakes answers that must be grounded and defensible.

For a lookup that maps to a single passage, one-shot RAG is faster, cheaper, and completely adequate. Do not grow a spine where a reflex will do. But the moment your users start asking real questions — the kind that span documents and require the system to reason about its own uncertainty — passive retrieval breaks, and the loop is what holds the answer up.

Ask a naive RAG system a question like “Which of our regions missed their SLA last quarter, and what caused it?” and watch it faceplant. It embeds the whole sentence, pulls the top eight chunks that look vaguely similar, stuffs them into a prompt, and hopes the model can reason...