Northwind Diagnostics
An AI receptionist that was measured before it was switched on
Northwind’s eleven collection centres shared a call queue that abandoned 23% of inbound calls at peak. A voice agent now handles 71% of them end to end — booking, rescheduling and results enquiries — and the calls that reach a person reach one faster than they used to.
- Client
- Northwind Diagnostics
- Sector
- Healthcare
- Run
- 5 months · Mar–Jul 2025
- Filed
- 2025-08-19
- Capability
- AI and generative AI solutions · Business automation · API and system integration · Cybersecurity
- Stack
- Azure OpenAI · Azure AI Speech · .NET 8 · ASP.NET Core · Azure AI Search · PostgreSQL · Twilio · Azure Container Apps · OpenTelemetry

Impact
71%
Inbound calls handled end to end
23% → 4%
Call abandonment rate at peak
380
Scored cases in the evaluation suite
Challenge
Healthcare
5 months · Mar–Jul 2025
Four receptionists covered eleven sites, and the queue was worst at exactly the times patients called: the first hour of the morning and the hour after work. Abandoned calls did not go away, they returned as walk-ins at a centre with no slot free. The clinical board had rejected an earlier chatbot proposal, and were right to: it had no measurement plan, no route to a human, and no answer to what happens when it gets a patient’s question wrong.
Solution
Nothing was switched on for the first four weeks. We instrumented the existing queue, recorded handling time, abandonment and reason-for-call across 4,100 calls, and built an evaluation set of 380 real transcripts with agreed correct outcomes — including the awkward ones, where the caller is distressed or asking for a result the agent must never read out. The agent was then built against that set and runs it in CI on every prompt change. Retrieval is scoped per caller after identity verification, so it cannot surface a record the caller could not otherwise obtain. Anything below the confidence floor, anything clinical, and anything the caller asks to escalate transfers to a person with the transcript attached. It ran on one site for three weeks before the other ten.
Plates

Plate 01Every branch below the confidence floor ends at a person. 
Plate 02The evaluation set runs on every prompt change, before deployment. 
Plate 03The four-week baseline is what makes the containment figure defensible.
Testimony
What convinced our clinical board was not the demo. It was that they measured four weeks of calls before switching anything on, so when they told us it was handling seven in ten we could check.