AI2026-06-30
What an AI pilot should cost you before it earns anything
Most AI pilots stall at the second budget review, and it is almost never because the model underperformed. It is because nobody built the thing that would have proved it did not.
- Author
- Mei-Lin Chua
- Published
- 30 JUN 2026
- Read
- 6 MIN
- Ref
- 242E2B
A generative AI pilot is easy to start and unusually hard to finish. The demo takes a fortnight and impresses everybody. Eight months later the same system is still described as a pilot, the enthusiasm has gone, and nobody can say precisely why.
The failure is almost never the model. It is that the pilot was built without the apparatus needed to evaluate it, so when the second budget conversation arrives there is nothing to put on the table except anecdotes — and anecdotes lose to a spreadsheet.
Here is what that apparatus costs, roughly, and why each part is not optional.
The baseline: two to four weeks
Before anything is switched on, record what the current process costs. Handling time, volume by category, error rate, escalation rate, abandonment — whatever the equivalents are for the process in question.
This is the single most common omission and the most expensive one. Without it, every claim about the system afterwards is relative to a number somebody remembered.
It also routinely changes the use case. On one engagement the baseline showed that the category the client wanted to automate was 6% of volume. The category they had not mentioned was 44%.
The evaluation set: one to two weeks, then forever
An evaluation set is a fixed collection of real cases with agreed correct outcomes. Not synthetic examples. Real ones, including the awkward cases, the ambiguous ones, and the ones where the correct behaviour is to refuse.
Two to three hundred cases is usually enough to be informative. Getting to that number takes a week or two of work with someone who knows the domain, and it is the least glamorous part of the whole project.
It is also the part that makes everything afterwards possible:
- A prompt change can be tested rather than argued about.
- A model upgrade can be evaluated in an afternoon instead of being feared.
- A regression in one category becomes visible when it happens, rather than three months later through complaints.
Run it in CI. If the evaluation set only runs when someone remembers, it will stop running by week six.
The escalation path: designed, not appended
Every agent needs a confidence floor and a route to a person, and the route has to carry the context. A hand-off that makes the customer start again is worse than no automation at all, because it adds a step to the process it was meant to shorten.
Budget for this properly. In our experience the escalation path is between a fifth and a third of the build, and it is the part that determines whether the operational team supports the project or quietly undermines it.
The safety work: one to two weeks
Prompt injection testing, permission scoping on retrieval, a decision about what the agent must never say, and a fallback for when the model provider has an outage. Providers do have outages, and a system that simply stops is a system that goes back to the manual process without warning anybody.
Permission scoping deserves specific attention. An assistant that retrieves across a document store will happily surface a document its user could not otherwise open, unless retrieval is filtered by that user's permissions at query time. This is a data breach with a friendly interface, and it is the finding we most often bring to clients who built the pilot themselves.
The measurement afterwards: ongoing, small
A dashboard covering containment, escalation rate and cost per interaction, and a monthly comparison against the baseline. Half a day a month, at most.
Model providers ship changes to their systems. A deployment nobody re-measures degrades silently, and the first signal is usually a complaint rather than a metric.
The shape of the estimate
Across the first deployments we have run, the split lands roughly here. The figures vary by use case; the ordering does not.
| Component | Share of build | Skipped when the pilot stalls |
|---|---|---|
| Baseline measurement | 10% | Almost always |
| Evaluation set and CI harness | 15% | Almost always |
| Model integration and prompting | 25% | Never — this is the part that gets demoed |
| Escalation path and hand-off | 20% | Frequently |
| Safety, permissions, provider fallback | 10% | Frequently |
| Integration with the system of record | 20% | Rarely |
Read the right-hand column as a list of the reasons pilots do not become deployments.
So what does it add up to?
For a typical first deployment, the model integration itself is a minority of the work — frequently under a third. The rest is baseline, evaluation, escalation, safety and integration with whatever system the output has to land in.
That ratio surprises people, and it is worth stating early, because a proposal that hides it produces a project that runs out of money at exactly the point where the interesting work starts.
The one thing to take away
If you are choosing between building the evaluation set and building two more features, build the evaluation set.
The features can be added later by anyone. The evaluation set is what lets you tell whether adding them helped — and, at the budget review, whether any of it was worth doing.