Skip to content

Tensions in Practice: Late retry recovery in tension with bounded memory

Repeated requests · lost replies

A service may finish a request just before its reply is lost. The caller retries with the same operation key so the service can recognize work it has already done. But that recognition depends on a retained record. Keeping it longer allows later recovery; expiring it requires an explicit rule for old requests, or a forgotten retry can look new.

Recover after a late retry

Recognize earlier work even when the caller returns after a long interruption.

Bound retained state

Limit the records that must remain available and correct for duplicate detection.

Why these aims pull against each other

The retry promise lasts only as long as the required identity can be recognized. Extending that promise costs retained state; enforcing its end makes some legitimate recovery attempts fail visibly.

Compare the arrangements

Retain longer

Retain operation identities and their outcomes for a longer documented recovery window.

What it protects
A retry within that window can recover the prior outcome without repeating the effect.
What it costs
More records must be stored and maintained; a longer finite window still ends.
When it fits
Late legitimate retries are important enough to justify the retained state, and identity/effect recording is correctly implemented.

Illustration note: The diagram assumes the original execution was recorded correctly and the retry uses the same key. Retention alone is not an exactly-once protocol.

Expire and reject

After the shorter documented window, reject a request carrying an older issuance time.

What it protects
Deduplication memory can be bounded without silently replaying expired operations.
What it costs
A legitimate late caller receives an error and needs a separate reconciliation path.
When it fits
The system can enforce trustworthy key age and callers understand expiry before relying on retry.

Illustration note: The visible error is source-supported. Reconciliation is an explicitly unresolved application responsibility, not another execution shown here.

What this illustration does—and does not—establish

Idempotence: Idempotency Window Expiry vs Long-Tail Retry Reality supplies extended retention and explicit rejection. The example isolates the memory boundary and makes no claim about a particular provider.

  • No window size or storage estimate is supplied.
  • State-level duplicate protection does not automatically cover notifications or other downstream effects.
  • A lost, evicted, or inconsistent identity record can defeat either intended contract; the diagrams assume correct operation within the window.

Source entries

Idempotence

Prime · Source of the tension

Idempotence: Idempotency Window Expiry vs Long-Tail Retry Reality supplies the conflict examined here.

Idempotency Window Expiry vs Long-Tail Retry Reality

The corrective is to either extend the window to match the worst-case observed retry latency (with cost in cache size), or to stamp idempotency-key requests with their issuance timestamp and reject any incoming retry whose key is older than the documented window (so the failure mode becomes a visible client-side error rather than an invisible server-side double-execution), or to make the underlying operation naturally idempotent (state-convergence design) so that even out-of-window re-execution is safe.

Read the source section

Structural Tensions and Failure Modes

Many operations trigger downstream side effects (notifications, emails, downstream message emission, audit log writes, billing events, webhook fires) that are not naturally idempotent even when the underlying state change is.

Read the source section