Idempotency before retries
Adding retries to a non-idempotent endpoint converts a visible failure into an invisible one.
- A flaky downstream is causing errors and someone has opened a PR adding retry logic.
- Synchronous HTTP or RPC between services you do not own end to end.
Problem
A downstream service fails intermittently and someone has opened a PR adding retries. The change is three lines, the reasoning is sound, and it sails through review. The trouble is that a retry converts a visible failure into an invisible one: when a request gets processed twice nothing alerts, the data is just wrong.
Context
Synchronous HTTP or RPC between services you do not own end to end. If you cannot read the other side’s code or do not know their deploy schedule, this applies to you.
Approach
First establish what happens when this call is repeated. If the answer is “I don’t know,” adding retries is off the table, because you would be multiplying a behaviour you have not characterised.
Then put idempotency on the callee, not the caller. Each request carries a key generated by the client, and the server records the keys it has acted on. When the same key arrives again the work is not repeated and the original result is returned. The key’s lifetime should comfortably exceed the retry window; I usually keep them for a few days.
It matters that the client generates the key. A server-generated identifier changes on retry and is useless for this.
Only once that is in place do I add retries, with exponential backoff and a bounded attempt count.
Tradeoffs
Storing idempotency keys costs storage and a lookup; on high-volume endpoints that cost is real. If the key record is not written in the same transaction as the work itself, it becomes a new source of inconsistency rather than a fix for one.
There is also added latency, since every request now performs one more check.
When this does not work
If the call is already naturally idempotent this is unnecessary. A write that sets a record to a final value, or any read, does not need it.
It is also insufficient when the side effect is external. On an endpoint that sends an email or moves money, an idempotency key prevents the second call but tells you nothing about whether the first one actually completed; that needs a separate reconciliation step.