Retries can turn a brief dependency slowdown into sustained overload. Each caller adds more work while the dependency is already struggling, and several service layers may multiply one user request into dozens of attempts. A timeout without a deadline budget has a related failure: lower layers keep working after the caller has stopped caring.
Design timeout and retry behavior together. Bound the total operation, retry only outcomes that are both transient and safe to repeat, spread attempts over time, and expose when the recovery mechanism is consuming capacity.
Start with one end-to-end deadline
An operation has a total time budget derived from its caller, user experience, queue lease, or service objective. Pass the remaining deadline through downstream calls instead of giving every layer a fresh full timeout.
Reserve time for work after a dependency response: parsing, persistence, cleanup, and returning to the caller. A five-second outer deadline with a five-second dependency timeout leaves no recovery margin.
Configure connection establishment, TLS handshake, response headers, response body, and idle stream behavior according to the client library's actual timeout semantics. A single “request timeout” option may not cover DNS or socket acquisition. Verify it under the named runtime rather than assuming the label covers every phase.
Cancellation should propagate where the protocol and library permit it. The server still needs idempotency and reconciliation because cancellation may arrive after the dependency has committed the effect.
Retry only when repetition is valid
A retry needs two independent decisions:
- Is the failure plausibly transient?
- Is repeating the operation safe?
Connection resets, selected timeouts, and explicit overload responses may be transient. Validation errors, authentication failures, and most authorization failures usually are not. Status code alone can be insufficient; API contracts should document retry behavior and Retry-After where appropriate.
Safe repetition comes from HTTP method semantics, an idempotency key, a transaction state machine, or a domain uniqueness constraint. Do not retry a non-idempotent write merely because the client did not receive a response.
Place retries at one layer when possible. If an HTTP client, service method, job worker, and gateway each attempt three times, the dependency can receive many calls for one original action. Let the layer with the best knowledge of deadline and idempotency own the policy.
Back off and add jitter
Immediate retries synchronize callers and increase pressure. Exponential backoff grows the delay between attempts, while a cap prevents unbounded waiting. Jitter randomizes timing so clients do not wake in a coordinated wave.
The AWS Builders' Library article on timeouts, retries, backoff, and jitter explains how retries are selfish from the server's perspective and why jitter helps spread retry traffic. The exact formula should be part of the client contract and tested statistically.
A generic policy needs:
- maximum attempts including the initial call;
- base delay and maximum delay;
- jitter strategy;
- total deadline;
- retryable outcome set;
- server-directed delay handling;
- cancellation behavior.
Do not sleep past the remaining deadline. Clamp a server-provided delay to policy only when the API contract permits; ignoring it can worsen overload, while obeying an unbounded value can strand work.
Protect the dependency and the caller
Limit concurrent attempts, not only attempts per operation. A retry budget can cap the fraction of traffic spent on retries. Queue limits and load shedding prevent recovery traffic from exhausting worker pools.
A circuit breaker can stop calls after a defined failure condition and probe for recovery later. It introduces state and can create sharp behavior changes, so it should not replace timeouts, concurrency bounds, or dependency capacity planning. Keep its state observable and partition it at a boundary that matches the failure domain.
Hedged requests, which send another copy before the first finishes, can reduce tail latency for idempotent reads but intentionally increase load. Use them only with evidence, strict budgets, and cancellation of losing attempts. They are unsuitable as a default retry strategy.
Design worker retries as durable state
Background jobs may retry over minutes or hours. Persist attempt count, next eligible time, last outcome category, and stable operation identity. A worker restart must not reset the policy and create infinite attempts.
After the bounded retry set is exhausted, move the job to an explicit terminal or manual-review state. Preserve enough sanitized context to diagnose it. A dead-letter queue is not resolution; it is a holding area that needs ownership, retention, replay controls, and alerts.
Replaying several failed jobs simultaneously can recreate the original outage. Rate-limit recovery and confirm the dependency is healthy before releasing a backlog.
Observe original calls and retries separately
Record attempt number, total elapsed time, remaining deadline, outcome category, selected delay, and dependency identity as bounded telemetry. Distinguish original request rate from retry rate. A stable user-facing success rate can hide a rapidly growing retry tax.
Useful alerts include retry-budget exhaustion, rising attempts per operation, deadlines expiring before an attempt begins, queues accumulating terminal failures, and a high rate of ambiguous write outcomes. Avoid logging credentials, request bodies, or unbounded exception text.
Trace context should connect attempts to the original operation while giving each attempt its own span. This shows whether latency came from service work, waiting, or retry delay.
Verify the policy under controlled failure
Test delayed connection, slow response body, reset connection, overload response, invalid input, authentication failure, ambiguous write, cancellation, and dependency recovery. Confirm the number and timing of attempts, not just the final result.
The correct retry system reduces the effect of brief transient faults without adding unbounded work. If the team cannot state the deadline, retry owner, safe-operation rule, attempt cap, and overload response, the implementation is not yet a resilience control.
Was this guide useful?
Your rating helps us prioritize clearer, more practical technical content.
Review diffs, run checks, and prepare the release in Arezgit.
Keep Git review, security scanning, API checks, database inspection, and release preparation together in one local desktop application.
Explore Arezgit