A common pattern: an organisation runs a dozen AI pilots over a year, most of them technically successful, and puts none of them into production. The instinct is to blame model capability or vendor choice. Neither is usually the cause.
Pilots are designed to prove feasibility. Production requires four other properties, and none of them are tested by a successful demo.
One: the unit cost was never modelled
A pilot runs on a sample. Production runs on the whole volume, with retries, with longer context as the use case matures, and with a quality bar that pushes teams toward larger models.
Cost per task at production volume is a number worth computing before the pilot rather than after. It frequently changes which use case is worth doing, and occasionally reveals that the manual process it replaces is cheaper.
Two: data readiness was assumed
Pilots use a curated extract, prepared by someone who knew which records were clean. Production reads the live source, including the records with missing fields, inconsistent formats, and the historical decisions nobody documented.
This is where the schedule usually goes. Assess data readiness for the shortlisted use cases specifically, before committing to a delivery date.
Three: there is no evaluation suite
A pilot is judged by people looking at outputs and agreeing they seem good. That does not scale, does not survive a model version change, and gives operations no way to tell a regression from a bad day.
A production system needs a task suite, a regression set, and a scoring method agreed with the business owner. Building it is neither glamorous nor expensive, and its absence blocks every deployment decision that follows.
If you cannot answer "did last week’s change make it better or worse" with a number, the system is not ready for production, whatever its demo looked like.
Four: nobody owns a wrong answer
The question that stops more pilots than any technical constraint: when the system is wrong, and it will be, who is accountable, what does the affected person experience, and what is the correction path?
This is a governance design problem, not a modelling one. It needs a named owner, a documented error taxonomy, a correction mechanism, and a threshold at which the system is withdrawn. Organisations that answer it before piloting move faster, because the question does not surface for the first time at the production review.
A better pilot design
- 01Model the production unit cost first, and drop the use case if the economics do not work at volume.
- 02Read the live data source during the pilot, not a curated extract.
- 03Build the evaluation suite before the pilot, and use it to judge the pilot.
- 04Name the accountable owner and the correction path before the first user sees an output.
- 05Set the go and no-go thresholds in advance, so the decision is not a negotiation after the fact.
A pilot run this way takes perhaps two weeks longer and produces a decision rather than a demo.