Skip to content

Engineering2026-03-17

The integration that works until it is retried

Idempotency is not an advanced topic. It is the difference between an integration that can be operated and one that quietly creates duplicate records for eighteen months.

Author
Farida Khan
Published
17 MAR 2026
Read
5 MIN
Ref
75F78C

Integration failures are rarely dramatic. Nothing goes down. The dashboard stays green. Instead, somewhere in a finance system, the same payment appears twice, and it is found six weeks later by a person doing a manual check.

The cause is almost always the same, and it is not a hard problem. It is a problem that gets deferred because the happy path works.

What actually happens

Your system receives a webhook. It processes it, writes a record, and returns a response. The response is lost — a timeout, a load balancer resetting a connection, a deployment mid-request. The sender, correctly, retries.

You now have two records for one event.

Every part of that sequence is normal behaviour. The sender is doing what its documentation says it will do. The network is doing what networks do. The only thing that is wrong is the assumption that a message arrives once.

The rule

Every message needs a key that identifies the event, not the delivery. Store the key with the record. Before processing, check whether the key has been seen; if it has, return the original result rather than doing the work again.

Three details determine whether this actually holds:

The key must come from the sender. A key you generate on receipt is a new key on every retry, which is the bug you were fixing. Most providers supply one — an event id, a message id, a delivery id. Use theirs.

The check and the write must be atomic. Two retries arriving simultaneously will both pass a check that is a separate query. A unique constraint on the key column, with the insert handling the violation, is the version that works under concurrency. A check-then-insert is the version that fails at exactly the moment volume makes it matter.

Return the original result, do not error. A retry is not a client mistake, and answering it with a failure sends the sender into a retry loop or a dead-letter queue for something that already succeeded.

Then: make the failures visible

Idempotency stops duplicates. It does not tell you when messages are not arriving at all, which is the other half of the problem.

The minimum that makes an integration operable:

  • A dead-letter queue, with an alert on depth greater than zero rather than a threshold. One stuck message is a signal.
  • A replay procedure an operator can run without a developer, tested at least once.
  • A daily reconciliation that counts both sides and reports the difference.

That last one is the deliverable nobody asks for and everybody needs. It is what turns "we think the integration is fine" into a number that arrives every morning. When it is not zero, you find out that day.

The three failure modes, and what each one needs:

FailureWhat you seeThe fix
Duplicate deliveryTwo records for one eventSender-supplied key, unique constraint, return the original result
Silent non-deliveryNothing at all, for weeksDead-letter queue with an alert on depth > 0
Partial processingOne side updated, the other notA daily reconciliation that counts both sides

Why this gets skipped

Because the happy path is a fortnight and this is another week, and the failure it prevents will not appear during the project. It appears in month four, in a different team's system, and by then it is eighteen months of history to unpick rather than a constraint to add.

The reconciliation report is the cheapest insurance in the whole build, and it is the first line struck from the estimate.

The test that costs nothing

Before you ship an integration, send the same message twice. Then send it fifty times in parallel.

If the result is one record, you are done. If it is two, or fifty, you have found the defect while it is still a code change rather than a data migration.

Farida Khan, Principal Architect

Written by

Farida Khan

Principal Architect

Designs the systems and, more usefully, the sequence that gets a client to them while the current platform stays up. She led the Portside Insurance modernisation across eighteen months without a feature freeze. Farida argues for the boring option roughly nine times in ten and is usually right; the tenth time is why she is worth arguing with.

Share

Related notes

All notes
  • 01Engineering

    Moving off .NET Framework without a rewrite

    Almost every .NET Framework application we are asked to replace should be migrated instead. Here is the order the work goes in, and the four things that actually block it.

    .NET · Migration6 MIN
  • 02Delivery

    Measure the process before you automate it

    The most expensive automation failures we are called in to fix all share one property: nobody recorded what the manual process actually cost, so nobody can tell whether the replacement is better.

    Automation · Measurement5 MIN
  • 03AI

    What an AI pilot should cost you before it earns anything

    Most AI pilots stall at the second budget review, and it is almost never because the model underperformed. It is because nobody built the thing that would have proved it did not.

    Generative AI · Evaluation6 MIN