← Insights · August 7, 2026

Why Most AI Pilots Never Reach Production

There is a well-worn statistic, quoted in various forms, that the large majority of AI pilots never make it into production. Whether the real number is sixty percent or eighty is beside the point. Anyone who has worked inside a mid-sized organisation over the last three years has seen it happen at least once: a promising demo, an enthusiastic steering group, and then a slow fade into the backlog.

What is interesting is that the reasons are remarkably consistent, and almost none of them are about the model.

Reason one: the pilot was never designed to be productionised

A pilot built to impress and a pilot built to become a product are different artefacts. The first runs on a laptop against a CSV export, has no authentication, no logging, and one hard-coded happy path. The second runs against a live data source, handles the awkward cases, and has an evaluation set attached to it.

The trouble is that the first one is faster to build and demonstrates just as well. So it gets built, it goes down brilliantly, and then someone asks how long it will take to roll out. The honest answer is that the pilot contributed almost nothing reusable and the real work starts now. That conversation kills projects, because the budget was set on the basis of the demo.

The fix is not to make pilots heavier. It is to be explicit at the outset about which kind you are building and to price the second phase before you build the first.

Reason two: nobody owns the outcome

Ask who will own the system a year after go-live. If the answer involves a pause, the project is in trouble. Sponsorship is not ownership. A sponsor approves budget; an owner deals with a user complaining that the classification was wrong, decides whether the threshold should change, and argues for the maintenance budget in next year planning round.

Systems without an owner do not fail dramatically. They degrade. The data source changes shape, quality drops, users lose confidence, usage falls, and eighteen months later someone asks whether we still need this and nobody defends it.

Reason three: the data was not really available

This is the most common technical blocker and it is almost always discovered late. The data exists, but it sits in a system whose vendor charges for API access. Or it is available but contains personal data that nobody has assessed a lawful basis for using this way. Or it is accessible but the field everyone assumed was reliable is only populated in forty percent of records.

A pilot can paper over all three by using a hand-curated extract. Production cannot. We now front-load a data assessment before any build precisely because this failure mode is expensive and predictable.

Reason four: no definition of good enough

If you cannot say what accuracy, coverage or handling time the system needs to hit, you cannot say whether it is ready. What happens instead is that the system is judged anecdotally. Someone senior tries three examples, one of them is wrong, and the project acquires a reputation it never recovers from.

An evaluation set of a few hundred labelled examples, agreed with the business before the build, changes that conversation entirely. You can point at a number, show performance by segment, and have an adult discussion about whether ninety-one percent with a human review path is better than the current manual process. Usually it is, by a wide margin, and nobody had ever measured the manual process either.

Reason five: the production requirements were invisible

Authentication, authorisation, audit logging, rate limiting, cost ceilings, monitoring, alerting, incident runbooks, data retention, backup, disaster recovery, penetration testing, accessibility, and a change process. None of these appear in a demo. All of them are required before a system touches real users, and together they are usually more work than the AI component.

This is not a reason to despair. It is a reason to budget honestly. In our experience the model or prompt work is somewhere between ten and twenty percent of the total effort on a production AI system. Teams that plan on that basis ship. Teams that assume the demo was eighty percent of the way there do not.

What a project that ships looks like

None of that is exotic. It is ordinary software engineering discipline applied to a category of work that has attracted an unusual amount of magical thinking.

If you have a stalled pilot

Most stalled pilots are recoverable, but the recovery usually involves rebuilding rather than extending, and it is better to know that early. We do a short, fixed-fee review that tells you what is salvageable, what the real production path costs, and whether the use case is worth it at that price.

Read more about our AI implementation services and our AI readiness assessment, or get in touch.

Related reading

How we can help

AI Applied is a UK software studio building production AI systems. See our AI services, including AI readiness assessment, AI implementation, machine learning development, data engineering, AI automation and AI governance and compliance. Or just get in touch.

Insights

Occasional, useful notes on applied AI.

What's actually working, what to ignore, and what the new regulation means for UK businesses. No spam.

We’ll only use your email address to send you these Insights notes. We never share it, and you can unsubscribe from any email. See our Privacy Policy.

AI Services

Where we work

Company

Latest writing

AI Applied Ltd, Technology House, 9 Newton Place, Glasgow G3 7PR. Registered in Scotland SC806963. support@aiapplied.uk ยท +44 141 465 5233