Saltar al contenido
Platform

89% of AI pilots never reach production

The reason is almost never the model. It is five conditions around it, and all five can be diagnosed before you sign anything.

Luis Aguilar 4 min read

There is a conversation that repeats in committees this year. Someone presents an AI agent pilot that went well, the team approves continuing, and twelve months later the project is still a pilot.

It is not bad luck or bad execution. It is a pattern, and it has been measured.

The size of the problem

AI agent pilots that never reach production by source, 2026
Deloitte tech trends 2026 89%
Independent studies range reported in 2026 86% to 88%
Only the two sources measuring the same thing. Gartner separately estimates that over 40% of agentic projects risk cancellation before 2027, but that measures something else and putting it on the same bar would make them look comparable. Deloitte Tech Trends 2026 · independent studies compiled in 2026

The figure that best frames the conversation is another one: 78% of companies have at least one pilot running, and only 14% have scaled one across the organisation. Almost the entire market sits at the same halfway point.

What fails is not the model

It is worth saying plainly because it saves months: the same model that runs in the pilot runs in production. If the pilot worked, the model is fine.

What changes is everything around the model. The five most repeated reasons in the 2026 analyses are always the same:

  1. Integration with systems that already exist. The pilot read a copy of the data. Production demands reading and writing in the real system, with its permissions, its latency and its edge cases.
  2. Inconsistent quality at volume. A hundred hand-picked cases behave differently from a hundred thousand real ones.
  3. No monitoring. If nobody sees what the agent decided and why, nobody can authorise it to continue.
  4. Nobody owns it. The pilot belonged to innovation or technology. The operation that should absorb it never took part, and does not claim it.
  5. Not enough domain data. The agent knows about everything and nothing about this institution.

None of those five is fixed by changing model. All five can be spotted before signing.

What separates a pilot from an operation
Criterio Pilot Operation
Where the data comes from A copy prepared for the trial. The real system, with its permissions and its latency.
Who supervises Someone reviews every case. Nobody watches. The limits are declared and auditable.
Who is accountable Innovation or technology. A named person inside the operation.
What it costs Calculated on the trial. Calculated at real volume.
A pilot proves something can work. An operation proves it works with nobody watching. The gap between them is not technical: it is who defines the limits and where they are written down.
CrediAmigoAgent · actionslive
buscarUsuarioDatabaseLook up the customer by ID
clasificarUsuarioLogicClassify the credit level
requisitosDataValidate approval requirements
aprobacionLogicConditions to approve credit
cuotaMensualComputeCompute the monthly payment
Code InterpreterWeb Search
In an operation, what the agent can and cannot do is declared and auditable: what it checks, how far it negotiates, when it hands the case to a person. A pilot usually works because someone supervises every case, and that does not scale. See the controls and traceability →

Shimli operates today with 160 active clients across 15 countries, with agents in production on collections, credit, support and internal operations. That is the difference this article tries to make visible: not a better model, but a process that already made it from trial to operation.

Questions on this topic

Why does a pilot that worked never reach production?
Because a pilot is assessed with hand-picked cases and constant supervision, and production demands the opposite: volume, edge cases, real data and nobody watching. The five most repeated reasons are integration with legacy systems, inconsistent quality at volume, lack of monitoring, no clear owner inside the organisation, and a shortage of domain data.
What sets apart a pilot that will scale?
Three things you can measure before signing: that the agent already reads and writes in the real systems and not a copy, that a named person inside the institution is accountable for the outcome, and that cost per task is calculated at volume rather than in the trial.

Put an agent to work on this process.

Create the account and start today. Or we take 30 minutes and define which metric you'll move.