Start here: the retry you can't avoid#
TL;DRthe 30-second version
- A timeout can't tell you whether the request worked. So the caller has to retry, and the server will sometimes see the same request twice.
- Idempotency means applying the same request twice has the same effect as applying it once.
- The pattern: the client attaches an idempotency key, a unique id for this one operation, and reuses it on every retry. The server keeps a dedup store of key to saved result, does the work once, and replays the saved result for repeats.
- Exactly-once delivery is impossible over a network. You build effectively-once processing instead: at-least-once delivery plus a receiver that ignores duplicates.
- The other half is retry discipline: exponential backoff plus jitter, so a wave of failures doesn't become a synchronized retry storm.
Your service sends a request to charge a customer $50 for an order. It waits. Nothing comes back, and after a few seconds the call times out. Now you have to decide what to do, and you're missing the one fact you need: did the charge happen?
A timeout can't tell you. Maybe your request never reached the payment service, and no money moved. Maybe it did reach the service, the charge succeeded, and only the reply got lost. From where you stand, both look the same: you sent a request and got silence.
You can't just give up. That would abandon a payment that may have been needed, every time the network hiccups. So the caller retries. But a blind retry of a charge that already went through charges the customer a second time. Retries are forced on you, and they're dangerous. The whole job is to make them safe.
Build it: the idempotency key#
The goal is to make the operation idempotent: applying the same request twice has the same effect as applying it once. Send the $50 charge one time or five times, and the customer is charged $50, once. Then retries become boring.
So the server needs to tell 'a retry of a charge I already did' apart from 'a new charge that happens to look similar'. You can't dedup on the contents of the request. Two different customers can each buy a $50 item one second apart. Same amount, same product, genuinely two charges.
The fix is to let the client name the operation. Before it sends the charge, the client generates an idempotency key: a unique id that stands for this one intended action. A common choice is a UUID, a random 128-bit id with essentially no chance of two clients picking the same one. The client attaches that key to the request and reuses the same key on every retry of that same charge. New charge, new key. Retry of a charge, same key.
Now the server can dedup. It keeps a dedup store: a table that maps each idempotency key to the result of the request that carried it. The logic on every incoming request is three lines:
- Look up the request's idempotency key in the dedup store.
- If the key is new: do the work (charge the $50), then save (key → the response you produced) before replying.
- If the key is already there: skip the work and reply with the saved response.
Walk the timeout through it. The first attempt arrives with key idem_9f8c. The server charges $50 and saves idem_9f8c → "charged, receipt #A1", but the reply is lost. The client times out and retries with the same key. The server looks it up, finds it, and replays "charged, receipt #A1" without touching the card. One charge, two requests. This is what Stripe exposes as the Idempotency-Key HTTP header: you put a key you chose in that header, and two requests with the same key produce one result.
One detail matters: save the response, not just a 'done' flag. The retry needs the same answer the original got, the same receipt number and the same charge id, not a bare acknowledgement.
Exactly-once is impossible; effectively-once is the goal#
It's tempting to wish for a better network, where every message arrives exactly once. Then you'd never need keys. But exactly-once delivery over an unreliable network is impossible, and this is a classic interview trap.
The sender has two honest options. It can send once and stop. If the acknowledgement is lost, the message may have been delivered zero times. That's at-most-once. Or it can keep resending until it gets an acknowledgement. Then the message is delivered one or more times, because the acknowledgement itself can be lost. That's at-least-once. There is no third setting, because the sender can never be certain a delivery landed. The classic name for this is the Two Generals Problem: two generals who can only communicate by a messenger who might be captured can never be sure the other has read their message.
Real systems can't afford to lose a payment, so they choose at-least-once and live with duplicates. The trick is to move the 'once' from delivery to processing. Delivery can happen many times. Processing happens once, because the receiver's dedup store throws the duplicates away. The name for this is effectively-once, or exactly-once processing. Kafka's 'exactly-once' does no magic on the wire. Under the hood it is retry plus dedup, behind a clean API.
What it costs: the dedup window#
The dedup store can't grow forever. You keep a key only long enough to catch the retries that matter, then delete it. That retention period is the dedup window. Stripe keeps idempotency keys for 24 hours. After that a repeat of an old key is treated as a brand-new request.
Too short a window, and a client that retries after it expired finds its key forgotten. The server does the work again, and the double charge is back. Too long, and the store swells with keys no one will resend. The rule: the window must be longer than the longest retry horizon any client will use.
PredictYour payments API takes 5 million charge attempts a day. Each key plus its saved response is about 400 bytes, kept for 24 hours. Roughly how big is the dedup store? And what breaks if you cut the window to 5 minutes to save space?
Hint: Multiply attempts × bytes for the size. Then: what happens to a retry whose key was deleted before it arrives?
Size: 5,000,000 × 400 bytes ≈ 2 GB. That's small. Dedup storage is cheap, which is why nobody shrinks the window to save space. What breaks at 5 minutes is correctness. Any retry that arrives more than 5 minutes after the original finds its key deleted, so the server treats it as a new charge and bills the customer again. A client backing off for minutes during an outage, or a request stuck in a queue, blows past 5 minutes easily. Size the window by the client's maximum retry horizon, and keep it generously long: hours to a day, never minutes.
The second cost is the subtle one. Two copies of the same key can arrive at the same time. The first attempt is slow, the client times out and retries, and now the original and the retry are both inside the server at once. If both read the dedup store, both see 'key not found', and both charge the card.
The other half: retry without a storm#
Idempotency makes a retry safe. It says nothing about how often you retry, and getting that wrong turns a small outage into a big one. The failure is called a retry storm, or thundering herd.
Picture a server that slows down for a moment. A thousand clients time out at nearly the same time. If every client retries immediately, the struggling server is hit by a thousand retries on top of normal load. It slows down more, times out more requests, and triggers more retries. A server that would have recovered in a second stays down because the herd won't stop.
The first fix is exponential backoff: wait longer before each retry. Try after 1 second, then 2, then 4, then 8. But backoff alone isn't enough. All thousand clients failed at the same moment, so they all wait 1 second and retry together, then all wait 2 seconds and retry together. The herd is still synchronized. You've only spaced out its stampedes.
The second fix is jitter: add a random amount to each wait. Instead of every client waiting exactly 2 seconds, each waits a random time between 0 and 2 seconds. Now the retries spread across the window instead of landing in a spike, and the recovering server sees a trickle it can serve. Backoff plus jitter is the standard policy, and it's what AWS's SDKs ship with. AWS's own write-up concludes that the jitter, not the backoff, is what breaks up the herd.
Two more controls cap the volume. A retry limit stops a single request after a few attempts, because infinite retries just feed the storm. And a circuit breaker stops retrying a dependency that is clearly down, so clients fail fast instead of piling on.
Three ways to get effectively-once#
| Approach | How it dedups | Best for |
|---|---|---|
| Natural idempotency (PUT / set-to-value) | The operation overwrites, so a repeat is a no-op. No store needed. | Writes that set a value rather than accumulate |
| Idempotency key + dedup store | Client tags the request; server saves key → result and replays it | Side-effecting POSTs: payments, orders, sending a message |
| Idempotent producer / transactions (Kafka) | The system stamps each message and dedups for you behind the API | Streaming pipelines that want effectively-once without hand-rolled keys |
These are the same idea at different layers. If the operation overwrites, you're already done. If it accumulates and has a side effect, you need a key, or a system that supplies one for you.
When you don't need an idempotency key
- Reads change nothing, so they're already idempotent. Under HTTP's rules (RFC 7231), GET, HEAD, PUT, and DELETE are defined as idempotent. POST is not, which is why POST is the method that needs a key.
- Set-to-a-value writes are idempotent by shape. 'Set the shipping address to X' gives the same result whether it runs once or three times. Increment-style operations ('charge $50', 'add one item', 'append a row') duplicate on retry and need a key.
- Operations with a natural unique id can dedup on that instead. 'Create order #1234' can use #1234 as its own key. Mint a fresh key only when the request has no identity of its own.
Reach for a key when an operation has a side effect that duplicates on retry and no natural unique id to dedup on. That's most money-moving and write-once operations, and almost no reads.
In the wild
- Stripe: an Idempotency-Key header (a string you generate, up to 255 characters, a UUID recommended) on any POST. Stripe stores the key and its response for 24 hours and replays it for repeats. Reusing a key with different parameters is rejected, so a key can't be accidentally reused for a different charge.
- Kafka's idempotent producer (enable.idempotence=true): each producer gets a producer id and stamps every message with a per-partition sequence number. The broker remembers the last number it accepted and drops any message it has already seen. That's a key and a dedup store built into the broker.
- AWS SDKs retry with exponential backoff and jitter built in, and many AWS APIs accept a client request token (their name for an idempotency key), so an automatic retry of launching an instance doesn't launch two.
Pitfalls & gotchas
- Generating a new key on each retry. Every retry then looks like a new operation and charges the card again. Generate the key once, before the first send. Generating it inside the retry loop is the classic bug.
- Crashing after the work but before saving the key. The retry finds no key and double-charges. Write the charge and the dedup record in one database transaction. If they can't share a transaction, record the key as in-flight first, do the work, then mark it done, with a recovery step for anything left in-flight.
- A plain read-then-write dedup check. Two simultaneous retries both see 'not found' and both do the work. Use an atomic first-writer-wins check: a unique constraint on the key column.
- A dedup window shorter than the client's retry horizon. During an outage a client backs off for a long time, and a late retry finds its key already deleted. The window must outlast the longest retry any client will make.
In an interview
Idempotency comes up whenever the design moves money, sends something, or writes across a network. The interviewer is checking whether you know retries are unavoidable, so duplicates are a design problem and not an edge case. Lead with that.
Why do you need idempotency at all?
Because a timeout can't tell you whether the request succeeded. The client must retry, so the server will see duplicates. Everything follows from that one fact.
How does the idempotency-key pattern work?
The client generates a unique key for the operation and reuses it on every retry. The server keeps a dedup store mapping key to saved response. First time it sees a key, it does the work and saves the response. On a repeat, it replays the saved response. Saving the response, not just a flag, is what lets the retry get the same answer.
Can't a message queue give you exactly-once delivery?
No. Exactly-once delivery over an unreliable network is impossible (the Two Generals Problem). What queues advertise is exactly-once processing: at-least-once delivery plus dedup at the receiver. The marketing says 'exactly-once'; the mechanism underneath is always retry plus dedup.
What are the two costs you'd bring up unprompted?
The dedup window must outlast the client's retry horizon (Stripe keeps keys for 24 hours). And simultaneous retries need an atomic first-writer-wins check, a unique constraint on the key, not a plain read-then-write.
What about the retry side?
Exponential backoff plus jitter to avoid a retry storm, a retry cap, and a circuit breaker. Safety (idempotency) and rate (backoff) are two separate problems, and a good answer covers both.
PredictAn interviewer asks: 'You're designing a notification service. An upstream event can be delivered to you more than once, and each event should send exactly one email. How do you guarantee that?' What's the strong answer?
Hint: You can't stop the duplicates arriving. Where do you make them harmless, and what stops two copies racing?
Say up front that you can't get exactly-once delivery from upstream, so you build effectively-once at your end. Use a stable key from the event itself, the event id the producer already assigns, so redeliveries share a key. On each event, atomically claim the key with a unique-constraint insert before sending, so two copies racing in can't both send. Send only if you won the claim. Size the dedup window to outlast how long upstream might redeliver. Then add backoff plus jitter on your own outbound sends, so a mail-provider blip doesn't become a storm.
References
- Stripe API — Idempotent requests — The Idempotency-Key header, the 24-hour window, and the saved-response replay.
- RFC 7231 §4.2 — Safe and Idempotent Methods — Which HTTP methods are idempotent (GET, PUT, DELETE) and why POST is not.
- Exponential Backoff And Jitter (AWS Architecture Blog) — Marc Brooker's write-up showing jitter, not just backoff, breaks up a retry storm.
- Kafka — Idempotent producer — Producer id + per-partition sequence numbers: a key and dedup store built into the broker.