Skip to content

AI2026-06-30

What an AI pilot should cost you before it earns anything

Most AI pilots stall at the second budget review, and it is almost never because the model underperformed. It is because nobody built the thing that would have proved it did not.

Author
Mei-Lin Chua
Published
30 JUN 2026
Read
6 MIN
Ref
242E2B

A generative AI pilot is easy to start and unusually hard to finish. The demo takes a fortnight and impresses everybody. Eight months later the same system is still described as a pilot, the enthusiasm has gone, and nobody can say precisely why.

The failure is almost never the model. It is that the pilot was built without the apparatus needed to evaluate it, so when the second budget conversation arrives there is nothing to put on the table except anecdotes — and anecdotes lose to a spreadsheet.

Here is what that apparatus costs, roughly, and why each part is not optional.

The baseline: two to four weeks

Before anything is switched on, record what the current process costs. Handling time, volume by category, error rate, escalation rate, abandonment — whatever the equivalents are for the process in question.

This is the single most common omission and the most expensive one. Without it, every claim about the system afterwards is relative to a number somebody remembered.

It also routinely changes the use case. On one engagement the baseline showed that the category the client wanted to automate was 6% of volume. The category they had not mentioned was 44%.

The evaluation set: one to two weeks, then forever

An evaluation set is a fixed collection of real cases with agreed correct outcomes. Not synthetic examples. Real ones, including the awkward cases, the ambiguous ones, and the ones where the correct behaviour is to refuse.

Two to three hundred cases is usually enough to be informative. Getting to that number takes a week or two of work with someone who knows the domain, and it is the least glamorous part of the whole project.

It is also the part that makes everything afterwards possible:

  • A prompt change can be tested rather than argued about.
  • A model upgrade can be evaluated in an afternoon instead of being feared.
  • A regression in one category becomes visible when it happens, rather than three months later through complaints.

Run it in CI. If the evaluation set only runs when someone remembers, it will stop running by week six.

The escalation path: designed, not appended

Every agent needs a confidence floor and a route to a person, and the route has to carry the context. A hand-off that makes the customer start again is worse than no automation at all, because it adds a step to the process it was meant to shorten.

Budget for this properly. In our experience the escalation path is between a fifth and a third of the build, and it is the part that determines whether the operational team supports the project or quietly undermines it.

The safety work: one to two weeks

Prompt injection testing, permission scoping on retrieval, a decision about what the agent must never say, and a fallback for when the model provider has an outage. Providers do have outages, and a system that simply stops is a system that goes back to the manual process without warning anybody.

Permission scoping deserves specific attention. An assistant that retrieves across a document store will happily surface a document its user could not otherwise open, unless retrieval is filtered by that user's permissions at query time. This is a data breach with a friendly interface, and it is the finding we most often bring to clients who built the pilot themselves.

The measurement afterwards: ongoing, small

A dashboard covering containment, escalation rate and cost per interaction, and a monthly comparison against the baseline. Half a day a month, at most.

Model providers ship changes to their systems. A deployment nobody re-measures degrades silently, and the first signal is usually a complaint rather than a metric.

The shape of the estimate

Across the first deployments we have run, the split lands roughly here. The figures vary by use case; the ordering does not.

ComponentShare of buildSkipped when the pilot stalls
Baseline measurement10%Almost always
Evaluation set and CI harness15%Almost always
Model integration and prompting25%Never — this is the part that gets demoed
Escalation path and hand-off20%Frequently
Safety, permissions, provider fallback10%Frequently
Integration with the system of record20%Rarely

Read the right-hand column as a list of the reasons pilots do not become deployments.

So what does it add up to?

For a typical first deployment, the model integration itself is a minority of the work — frequently under a third. The rest is baseline, evaluation, escalation, safety and integration with whatever system the output has to land in.

That ratio surprises people, and it is worth stating early, because a proposal that hides it produces a project that runs out of money at exactly the point where the interesting work starts.

The one thing to take away

If you are choosing between building the evaluation set and building two more features, build the evaluation set.

The features can be added later by anyone. The evaluation set is what lets you tell whether adding them helped — and, at the budget review, whether any of it was worth doing.

Mei-Lin Chua, Head of AI Engineering

Written by

Mei-Lin Chua

Head of AI Engineering

Builds the assistants and voice agents, and — the part that matters — the evaluation sets that decide whether they are working. She built the Northwind Diagnostics receptionist and insisted on four weeks of baseline measurement before a single call was answered by it, which is why the deflection figure can be defended.

Share

Related notes

All notes
  • 01Delivery

    Measure the process before you automate it

    The most expensive automation failures we are called in to fix all share one property: nobody recorded what the manual process actually cost, so nobody can tell whether the replacement is better.

    Automation · Measurement5 MIN
  • 02Engineering

    Moving off .NET Framework without a rewrite

    Almost every .NET Framework application we are asked to replace should be migrated instead. Here is the order the work goes in, and the four things that actually block it.

    .NET · Migration6 MIN
  • 03Cloud

    The restore you never tested is not a backup

    Every organisation we assess has backups. Roughly half have never restored from them. The gap between those two facts is where recovery plans go to die.

    Disaster recovery · Azure5 MIN