Last month a customer walked me through a slide deck with eleven “agents” on it. By the end of the hour we had found one real agent, three workflows, two chatbots, and five ideas that nobody had scoped yet. Nobody in that room was being dishonest. The word “agent” has simply stretched to cover almost anything with a model behind it.
That is the real job of a principal right now. Customers do not need another person cheering for agents. They need someone who can sort what is viable from what is not, then help them build the right thing, which is sometimes an agent and often is not.
Why This Matters Now
The gap between buzz and production is wide. A Gartner 2026 CIO survey found that only 17% of organizations have deployed agents, while more than 60% expect to within two years. The Sinequa State of Enterprise Agentic AI 2026 report is even more sobering: only 24% have deployed a true agent, and only 10% have true agentic capabilities. Gartner also forecasts that more than 40% of agentic AI projects will be canceled by the end of 2027 because of rising cost, unclear value, and weak risk controls.
Notice what is missing from that list: model quality. In my experience, most agent projects do not fail because the model is weak. They fail because of how they are governed and run. The value was never clear, the costs crept up, the risk controls were thin, and nobody was assigned to manage the agent once it went live.
So I use a filter. Five questions, asked early, before anyone writes a line of orchestration code.
Filter 1: Is It Really an Agent?
The industry now has a name for the relabeling problem: agent washing. Chatbots, assistants, and RPA tools are being renamed as agents, and buyers are paying agent prices for them.
My test is simple. A real agent does three things:
- Plans. It breaks a goal into steps it was not explicitly given.
- Uses tools. It calls APIs, queries data, or takes actions in other systems.
- Acts on its own. It decides the next step based on what it just observed, without a human or a fixed script choosing for it.
If the steps are known in advance, you have a workflow. Build it as a workflow. A state machine with a model call in one or two steps is cheaper, easier to test, and easier to explain to an auditor. I tell customers this is not a downgrade. It is the correct design.
Fixed steps, known branches -> Workflow (Step Functions, plus a model call)
Answer questions, no actions -> Assistant / RAG
Repetitive UI clicks -> RPA
Open ended goal, tool choice -> Agent
Filter 2: What Happens When It Fails?
Every agent demo works. That is what demos are for. The question is what happens on day thirty when the input looks different. Agents that shine in demos break on new input formats, loop on vague instructions, or quietly make up data. The last one is the dangerous one, because a confident wrong answer does not trigger an alarm.
I ask the team to describe three failure paths out loud:
- Catching it. How do you know the agent failed? Schema validation on outputs, checks against a system of record, and confidence thresholds that route to a human.
- Containing it. What is the blast radius? Scoped credentials, read only access by default, step and spend limits per task, and an approval gate before any irreversible action.
- Recovering from it. Can you replay the trace, see which tool call went wrong, and roll back what the agent did?
If the answer to any of these is “the model is pretty good, so it should be fine,” the project is not ready.
Filter 3: Can You Measure and Test It?
You cannot improve what you cannot measure, and you cannot ship what you cannot test. Before production, I want to see an evaluation suite built on the customer’s real edge cases, not the vendor’s happy path examples.
A useful starting point is fifty to two hundred real tasks pulled from tickets, logs, or past work, each with a known good outcome. Run the agent against them on every prompt change, model change, and tool change. Something as plain as this works:
results = []
for case in eval_cases:
outcome = agent.run(case["input"], max_steps=15)
results.append({
"id": case["id"],
"passed": grade(outcome, case["expected"]),
"steps": outcome.step_count,
"cost_usd": outcome.cost_usd,
})
pass_rate = sum(r["passed"] for r in results) / len(results)
print(f"pass rate: {pass_rate:.1%}")
Then carry the same idea into production. Sample live traces, grade them, and track pass rate, step count, and escalation rate over time. Offline evaluation tells you whether to ship. Production monitoring tells you when to stop.
Filter 4: What Does Each Completed Task Cost?
Token pricing is the wrong unit. Customers get excited about a fraction of a cent per call, then discover that one finished task took nine calls, two retries, a reasoning loop, and fifteen minutes of a person checking the result.
The number I ask for is cost per completed task:
cost per completed task =
(model calls + tool calls + retries + loops
+ infrastructure + human review time)
/ tasks completed correctly
The denominator matters as much as the numerator. If the agent completes 70% of tasks correctly, the other 30% still cost money and still need a person to finish them. Compare the result to what the work costs today. If the agent is not clearly cheaper, faster, or better at the same quality bar, the value case is not there yet. This is exactly the “rising cost, unclear value” pattern behind those cancellation forecasts.
Filter 5: Who Manages It Once It Is Live?
This is the question that kills the most projects, and the one teams skip most often. An agent is not a feature you ship and forget. It behaves more like a new team member who needs a manager.
Someone has to:
- Define the work. What tasks are in scope, and which are not.
- Set the quality bar. What counts as done, and what error rate is acceptable.
- Step in when things go wrong. Review escalations, pause the agent, and fix the prompt, tools, or data.
- Own the evaluation suite. Add new edge cases as production reveals them.
I ask for a name, not a team. If no single person owns that job, with time actually allocated to it, the project is not viable, regardless of how good the demo looked.
Practical Takeaways
- Classify before you build. Label every proposed “agent” as a workflow, assistant, RPA, or agent. Most will not be agents.
- Write the failure paths first. Catch, contain, and recover, documented before the first sprint.
- Build the evaluation suite from real data. Use the customer’s edge cases and rerun it on every change.
- Price the finished task, not the token. Include retries, loops, and human review.
- Name the owner. No operational owner, no production launch.
Saying “Not Yet” Is the Job
Principals earn trust by being the person in the room willing to say “not yet,” or “a simpler build will work better.” Customers remember who saved them from an expensive cancellation far longer than who sold them the shiniest demo.
And when the filter says an agent really is the answer, the right primitives exist. As I wrote in Agents Aren’t Web Requests, agents need runtimes built for long running, stateful, isolated work. Get the governance right first, and the infrastructure part becomes the easy part.
What questions are in your filter? I would like to hear what you ask customers before you let an agent near production.