Engineering2026-03-17
The integration that works until it is retried
Idempotency is not an advanced topic. It is the difference between an integration that can be operated and one that quietly creates duplicate records for eighteen months.
- Author
- Farida Khan
- Published
- 17 MAR 2026
- Read
- 5 MIN
- Ref
- 75F78C
Integration failures are rarely dramatic. Nothing goes down. The dashboard stays green. Instead, somewhere in a finance system, the same payment appears twice, and it is found six weeks later by a person doing a manual check.
The cause is almost always the same, and it is not a hard problem. It is a problem that gets deferred because the happy path works.
What actually happens
Your system receives a webhook. It processes it, writes a record, and returns a response. The response is lost — a timeout, a load balancer resetting a connection, a deployment mid-request. The sender, correctly, retries.
You now have two records for one event.
Every part of that sequence is normal behaviour. The sender is doing what its documentation says it will do. The network is doing what networks do. The only thing that is wrong is the assumption that a message arrives once.
The rule
Every message needs a key that identifies the event, not the delivery. Store the key with the record. Before processing, check whether the key has been seen; if it has, return the original result rather than doing the work again.
Three details determine whether this actually holds:
The key must come from the sender. A key you generate on receipt is a new key on every retry, which is the bug you were fixing. Most providers supply one — an event id, a message id, a delivery id. Use theirs.
The check and the write must be atomic. Two retries arriving simultaneously will both pass a check that is a separate query. A unique constraint on the key column, with the insert handling the violation, is the version that works under concurrency. A check-then-insert is the version that fails at exactly the moment volume makes it matter.
Return the original result, do not error. A retry is not a client mistake, and answering it with a failure sends the sender into a retry loop or a dead-letter queue for something that already succeeded.
Then: make the failures visible
Idempotency stops duplicates. It does not tell you when messages are not arriving at all, which is the other half of the problem.
The minimum that makes an integration operable:
- A dead-letter queue, with an alert on depth greater than zero rather than a threshold. One stuck message is a signal.
- A replay procedure an operator can run without a developer, tested at least once.
- A daily reconciliation that counts both sides and reports the difference.
That last one is the deliverable nobody asks for and everybody needs. It is what turns "we think the integration is fine" into a number that arrives every morning. When it is not zero, you find out that day.
The three failure modes, and what each one needs:
| Failure | What you see | The fix |
|---|---|---|
| Duplicate delivery | Two records for one event | Sender-supplied key, unique constraint, return the original result |
| Silent non-delivery | Nothing at all, for weeks | Dead-letter queue with an alert on depth > 0 |
| Partial processing | One side updated, the other not | A daily reconciliation that counts both sides |
Why this gets skipped
Because the happy path is a fortnight and this is another week, and the failure it prevents will not appear during the project. It appears in month four, in a different team's system, and by then it is eighteen months of history to unpick rather than a constraint to add.
The reconciliation report is the cheapest insurance in the whole build, and it is the first line struck from the estimate.
The test that costs nothing
Before you ship an integration, send the same message twice. Then send it fifty times in parallel.
If the result is one record, you are done. If it is two, or fifty, you have found the defect while it is still a code change rather than a data migration.