Back to Insights

Why AI Pilots Fail to Reach Production

The pilot proved the model works. It proved almost nothing about whether your organisation will let it run. The gap is a set of questions nobody was assigned to ask.

A pilot proves the model works. It proves almost nothing about whether your organisation will let it run, and that is the most expensive distinction missed in enterprise AI.

The gap between a pilot and production isn't a build. It's a set of questions nobody was assigned to ask.


The pilot was built to succeed

That was the brief. A POC runs on clean sample data, in a sandbox, with a hand-picked team, against criteria the vendor helped write. Nobody is being dishonest about it. The exercise was meant to find out whether the thing could work at all, and it can.

Better teams know this and insist on real data. They still get the easy half of it. Edge cases stay out, and there are always more of them than anyone expects. They aren't only in the data, either. They're in the process: what happens when five of them land in the same hour, who owns the exception, where it goes when nobody does.

Production runs on live pipelines at real volume. It touches customer data under actual privacy constraints. It needs an owner after go-live, an escalation path for when it errs, and someone willing to put their name to its outputs. None of that was tested, and none of it is technical.


The pilot answered "can this work". The question on the table is "should we run this"

And "should we" isn't one question. It's five, and the model working answers none of them.

  • Will we be able to.
  • Will it get approved.
  • Will the teams use it.
  • Will the resources still be there in nine months.
  • Will anyone change how they work because of it.

Companies make the production call on the evidence the pilot produced and call it de-risked, because something worked once. It isn't de-risked. It's unexamined. What follows are the three places that examination most often comes up short.


Who owns it once it is live

Most enterprise AI does one thing. It makes something legible that wasn't legible before. Where the delays sit. Which accounts get neglected. How long the approval really takes. Which of two teams is carrying the work.

That's the value, and it's also the threat. During a pilot the tension stays theoretical, because the numbers come from a sample, the scope is small, and nobody's quarterly review depends on them. At production scale the same output lands on a real person's record.

The person who's exposed rarely objects on those grounds. The objection arrives dressed as a data quality concern, or a compliance question, or a request to wait for the next planning cycle. Each one is reasonable on its own. Together they run out the clock.

The instinct is to handle it personally. Win them over, escalate, get air cover. That buys you one objection at a time, because the person was never the problem. Anything that changes how work happens crosses several teams, and control over the pieces is split across all of them. Any one of them can stall it, and any one of them, looking only at its own piece, has more to lose than to gain by going ahead.

Split the control and nobody in the chain is weighing a trade-off. They're each looking at a pure cost. The trade-off only exists for someone who owns the whole thing: accountable for the outcome, and holding enough control over the parts to make the call and carry it.

Without that person the programme has no decision-maker. It has a queue of vetoes. Decide who owns it while the project is still on paper, because after that the org chart has already answered.


What one error actually costs

99.9% accuracy sounds finished. It's the number that ends the meeting, and it shouldn't be, because an error rate on its own tells you almost nothing.

Run the system a hundred times a day and that one in a thousand arrives about three times a month. That's the arithmetic people do. Here's the one they skip: the wins and the losses aren't the same size.

Say the system saves a second per case. A thousand cases is roughly seventeen minutes saved. Now one of those thousand is wrong. Someone has to notice it, work out what happened, and repair what it touched downstream. If that takes seventeen minutes, you worked for nothing.

Errors cost more than successes save, because a success is silent and an error has to be found, understood, undone, and explained to whoever it reached. And usually you can't tell which one it was. If finding the bad output means checking all thousand, then checking is the new job, and the saving is gone before the error has cost you anything at all.

Then there's which case fails, and that isn't a random draw. The one that breaks is the unusual one, and unusual tends to mean consequential. The thousand you got right were the easy thousand.

So price the error before you buy the accuracy number. What does one bad output cost to catch, to fix, and to explain? Multiply that by three times a month. It's the figure the business case needed, and rarely the one on the slide.


Whether anyone is allowed to say no

Most AI guidance arrives from someone who benefits from the answer. The vendor who ran the pilot recommends production. The integrator who'd build it agrees. Everyone in the room is paid by the decision going one way.

An assessment that can only say yes isn't an assessment. An organisation that can't say no to an AI project doesn't have a process, it has a budget.


What a real gate looks like

We run the AI Readiness Assessment as a production gate. Three verdicts, and all three are real.

  • Go. The readiness is there and the remaining risk is ordinary. It is the shortest of the three to write.
  • Conditional Go. Named gaps, each carrying a price, a sequence and an owner. The word that matters is costed: a gap with no number and no name against it is an observation, and observations do not change what a steering group decides.
  • No-Go. The evidence does not support the spend.

No-Go is the verdict nobody commissions an assessment hoping for, and it's the most valuable one available. It costs a fee and saves a production budget, and it stops a programme that was going to fail on something nobody had been assigned to check.

That only works when the verdict isn't the sale. The fee is fixed and paid whichever way it lands, so a no-go costs us nothing to write. Whether we help with what comes next is a separate conversation, held after the verdict and never folded into it.


If your programme has a pilot behind it and no answer to who owns it, what an error costs, or who is allowed to stop it, the production decision is still open. It just looks closed.

Related: AI Readiness Assessment

Five pillars, one verdict, before the production budget is committed.

AI Readiness Assessment

Also relevant: The Briefing

One session that aligns a leadership team on what it is actually deciding.

The Briefing

Read next