Services Case Studies Insights About Start a project →

Why 88% of AI agent pilots never reach production.

AI Published July 3, 2026 9 min read

88% of AI agent pilots never reach production. The blockers are almost never the model: they are evaluation gaps, governance friction, and an identity and access model nobody budgeted for.

The gap between the demo and the production system

Every enterprise we talk to right now has an AI agent pilot running somewhere. Almost none of them have that pilot in production doing real work against real systems with real consequences if it gets something wrong. That gap is not a rounding error. According to 2026 survey data from Anaconda and Forrester, later replicated independently by a16z and an MIT Sloan CIO panel, 88 percent of agent pilots never make the crossing to production. S&P Global Market Intelligence and McKinsey put the number of enterprises with at least one agent actually live in production at 31 percent, and that figure is not evenly spread. Banking and insurance are out front at 47 percent, while healthcare and government trail at 18 percent and 14 percent respectively, which tells you something useful on its own: the industries with the clearest regulatory guardrails and the most mature identity and audit infrastructure are the ones getting agents into production, not the ones with the flashiest use cases.

This matters right now because the conversation in most boardrooms has quietly shifted. A year ago the question was whether agents could do useful work at all. Today the question is why a pilot that clearly worked in a demo still is not allowed to touch a production system, and the honest answer is rarely about the model. It is about everything that has to exist around the model before an organization will trust it with autonomy. The same three blockers recur in almost the same order every time.

Why most pilots stall, in order

When Forrester and Anaconda asked enterprise leaders what actually kept a working pilot from graduating, the answers clustered around three things: evaluation gaps, cited by 64 percent of leaders, governance friction, cited by 57 percent, and model reliability itself, cited by 51 percent. Notice what is missing from that list. Nobody is saying the agent could not perform the task. They are saying nobody could prove it would keep performing the task reliably, nobody could get the security and compliance sign-off needed to let it run unsupervised, and nobody had a way to catch a regression before a customer or a regulator did.

Evaluation gaps

A demo succeeds once, on a curated set of inputs, in front of an audience with every incentive to see it succeed. Production is the opposite: an open-ended stream of inputs nobody hand-picked, running continuously, with no audience and no forgiveness for a quiet regression that nobody notices for three weeks. Teams that treat evaluation as a one-time gate before launch are, in practice, deploying blind the moment anything upstream changes, whether that is a new model version, a prompt tweak, a tool schema change, or just a shift in how users phrase requests over time.

Governance friction

This is the security and compliance conversation that most engineering teams do not budget time for, and it is rarely about whether the agent is capable. It is about whether anyone can answer, on demand, who this agent is acting as, what data it can touch, and what happens if it does something wrong at 3am with nobody watching. Those are reasonable questions for a security team to ask before granting an agent access to a production database or a customer-facing action, and most pilots simply have no answer prepared.

Model reliability

This one is real but smaller than it looks in isolation. Reliability problems that would sink a pilot on their own often turn out to be scope problems in disguise. An agent asked to do one well-defined thing, with a narrow tool surface and a clear success condition, is dramatically more reliable than the same underlying model asked to handle an open-ended range of tasks. Which leads directly to the pattern that separates the agents that make it from the ones that do not.

Scope is the variable that actually predicts survival

The single clearest signal across every agent deployment we have reviewed is that narrow, tightly scoped agents graduate to production and broad, general-purpose ones do not, almost regardless of how capable the underlying model is. An agent that reconciles a specific category of vendor invoices against purchase orders, with a fixed set of tools and a well-defined definition of done, is something a team can evaluate exhaustively, monitor meaningfully, and roll back safely if it starts drifting. An agent pitched as a general finance assistant that can "handle anything in the AP workflow" is something nobody can fully test, because the input space is effectively unbounded, and unbounded input spaces produce unbounded failure surfaces.

This is a hard trade-off to accept, because the broad, ambitious version of the agent is the one that got funded and looks impressive in a steering committee deck. But the organizations actually shipping agents into production have, almost without exception, made peace with starting narrow. They pick one workflow, one data boundary, and one measurable outcome, get it reliable and observable, and only then expand scope, one bounded increment at a time. Skipping that sequencing is the most common reason a technically sound pilot gets stuck in an indefinite governance review.

The identity and access problem nobody budgets for

Here is the part that catches almost every engineering team off guard: an AI agent is a non-human identity that authenticates and takes action without a person directly in the loop for each step, and most access control systems were never built with that in mind. If your agent shares a service account with three other systems, or worse, runs under a developer's personal credentials because that was the fastest way to get the pilot working, you have no way to answer the questions a security review will ask. Who is this acting as. What can it touch. What happened the one time it touched something it should not have.

Production-grade agent deployments treat every agent as a first-class identity with its own workload identity or service account, scoped permissions that map to exactly what that agent's task requires and nothing more, and an audit trail that captures every action the agent took and the reasoning trace behind it. This is not a nice-to-have layered on afterward. It is frequently the difference between a security team approving a production rollout in weeks versus blocking it indefinitely, because without a named, scoped, auditable identity, there is no way to do an access review, no clean way to offboard the agent if it needs to be retired, and no way to contain the blast radius if something goes wrong.

  • Give every agent its own identity. Never share a service account across agents or reuse human credentials, even temporarily during a pilot, because that shortcut becomes permanent the moment the pilot works.
  • Scope permissions to the task, not the team. An invoice reconciliation agent needs read access to purchase orders and write access to a specific approval queue, not broad access to the finance system because it was easier to provision that way.
  • Log the reasoning, not just the action. When something goes wrong, "the agent updated this record" is far less useful during an incident review than the trace of why it decided to.
  • Design the offboarding path before launch. If you cannot describe how to revoke an agent's access cleanly, you are not ready to grant it in the first place.

Observability and evaluation as a running practice, not a launch gate

The teams that get past pilot stage have generally converged on treating observability as table stakes from day one, and an OpenTelemetry-first posture has become the de facto standard because it is vendor-neutral and portable across whichever agent framework a team ends up using or replacing later. But instrumentation alone is not the point. The point is the feedback loop it enables: observability surfaces the failure modes that happen in the wild, evaluation suites capture those specific cases as regression tests, and policy or prompt updates prevent them from recurring, in a cycle that never really finishes.

Evaluation itself needs to happen at more than one level of granularity to be useful. Trace-level evaluation checks a single full interaction end to end, run-level evaluation checks a batch of interactions against aggregate quality bars, and thread-level evaluation checks whether a multi-turn conversation stayed coherent and on task across its full length. Teams that only ever look at trace-level evals miss the slow, compounding drift that shows up over a long agent session, which is exactly the kind of failure that a demo, by its very short-lived nature, will never surface.

One number that surprises most engineering leaders the first time they hear it: infrastructure spend for a production agent deployment, covering observability, identity management, and the integration layer connecting the agent to real systems, tends to run 40 to 60 percent higher than the cost of the agent's core logic itself. Budgeting for the agent and forgetting everything that has to surround it is one of the most avoidable reasons a pilot's business case falls apart once someone prices out production.

What the twelve percent do differently

Across the pilots that do make it to production, the operating profile is unusually consistent, and none of the four traits are exotic or expensive on their own. There is named ownership, meaning one person or team is accountable for the agent's behavior in production, not a committee. There are scoped, measurable success criteria defined before launch, not adjusted afterward to match whatever the agent happened to do. There is automated evaluation running continuously against production traffic, not a one-time test suite that nobody revisits. And there is an organizational willingness to ship a narrow version, watch it, and roll it back without treating either the shipping or the rollback as a verdict on the whole program.

That last trait is the one that is hardest to manufacture after the fact, because it is cultural rather than technical. Teams that panic at the first rollback, or that treat a single agent failure as proof the whole initiative was a mistake, tend to freeze their pilots in permanent review rather than let them mature through the ordinary cycle of shipping something narrow, learning from what breaks, and expanding scope deliberately. If you are trying to get an agent pilot unstuck, the identity model, the evaluation pipeline, and the scoped rollout plan are the concrete things to build, and they are exactly the kind of production-readiness work our AI consultancy engagements are built around, but the willingness to treat a rollback as information rather than failure is the one thing no framework can hand you. Build for that from the start and the rest of this list gets considerably easier to execute.

Keep reading

Trying to get an agent pilot unstuck?

Start a conversation →
KT Solutions Assistant

Before we start, please share a few details so we can follow up with you.

Please enter your name and a valid email address.

End this conversation? Your chat will be emailed to us.