Skip to content

Cloud2026-04-22

The restore you never tested is not a backup

Every organisation we assess has backups. Roughly half have never restored from them. The gap between those two facts is where recovery plans go to die.

Author
Rohan Bhatt
Published
22 APR 2026
Read
5 MIN
Ref
72E0AB

The question we ask in every cloud assessment is not "do you have backups". Everyone has backups. The question is "when did you last restore from one, and how long did it take".

The answer is usually a pause.

Why the pause happens

Backups are configured once, by someone competent, and then they run. The dashboard is green. The retention policy is documented. Every part of the arrangement looks finished.

Restoring is different work. It is done under pressure, by whoever is available, against a system that is already broken, using a procedure that has typically never been executed. And the failure modes only appear at that moment:

  • The backup restores, but into a network that no longer exists in that shape.
  • The database restores, but the application cannot start because a secret was in a key vault that was in the same resource group.
  • Everything restores correctly, and it takes eleven hours against a recovery objective of two.

That last one is the most common and the most damaging, because nothing is technically broken. The plan was simply never timed.

What a drill actually involves

Half a day, once or twice a year, per critical workload.

Pick the objective first. Two numbers per workload: how much data you can afford to lose (recovery point) and how long you can afford to be down (recovery time). Get them from the business, not from IT. They are usually tighter than the technology currently supports, and finding that out is the point.

Restore into a clean environment. Not the existing one. A rebuilt environment tests the infrastructure definition at the same time, and if the infrastructure cannot be rebuilt from code, you have found a second problem worth knowing about.

Time it end to end. From "we have decided to invoke recovery" to "a user can log in and do their job". Not from the restore command to the restore completing. The gap between those two measurements is where most of the surprise lives.

Have someone who has not done it before run it. The procedure exists so it can be followed by whoever is on call at 4am, and the person who wrote it is not a fair test of whether it can be.

Write down the result with a date. Including the failures. A drill that surfaces three problems has done its job; one that surfaces none should be treated with suspicion.

What we find

The most frequent findings, in rough order:

  1. Secrets and certificates not covered by the backup, so the restored system cannot authenticate to anything.
  2. Recovery time three to five times the stated objective, mostly from steps nobody had timed.
  3. A dependency on a system that was assumed to be available and would not be, in the scenario that triggered the recovery.
  4. The procedure referring to a person by name who left in 2023.

None of these are exotic. All of them are invisible until somebody tries.

For reference, the last six drills we ran, with the stated objective against what the drill actually measured:

WorkloadStated RTOMeasuredGap found
Policy administration4h11h 20mCertificates outside the backup scope
Reporting warehouse24h6hNone — objective was conservative
Customer portal2h2h 40mManual DNS step, since automated
Document store8h31hRestore throughput never measured at volume
Identity1h1h 10mNone
Integration middleware4hFailedDepended on a system unavailable in the scenario

Two of the six met their objective. All six now have a dated result.

The cost of not doing it

An untested backup is not a technical position, it is a belief. It appears in board papers as a control, in risk registers as a mitigation, and in insurance documentation as a fact — and none of those consumers of the information can distinguish between a backup that would work and one that would not.

Half a day, twice a year, converts the belief into evidence with a date on it. It is the cheapest risk work available to most organisations, and it is skipped almost universally, because it produces nothing visible when it succeeds.

That is precisely why it needs to be scheduled rather than intended.

Rohan Bhatt, Head of Cloud & Platform

Written by

Rohan Bhatt

Head of Cloud & Platform

Runs the Azure practice and the delivery pipelines underneath it. He took Meridian Freight from a monthly release weekend to eleven deployments a week with the change failure rate going down rather than up. He will not sell you Kubernetes, and he will explain at length why not.

Share

Related notes

All notes
  • 01Delivery

    Why we will not take your pager permanently

    We run systems during delivery and for an agreed period after it. We will not do it indefinitely, and the reason is not capacity — it is that permanently outsourced on-call makes software worse.

    Operations4 MIN
  • 02Delivery

    Measure the process before you automate it

    The most expensive automation failures we are called in to fix all share one property: nobody recorded what the manual process actually cost, so nobody can tell whether the replacement is better.

    Automation · Measurement5 MIN
  • 03AI

    What an AI pilot should cost you before it earns anything

    Most AI pilots stall at the second budget review, and it is almost never because the model underperformed. It is because nobody built the thing that would have proved it did not.

    Generative AI · Evaluation6 MIN