Resilience · Anti-pattern
Retry storm
When several layers each retry failed calls, one request turns into dozens of attempts, all aimed at the part that is already failing.
GitHub had an outage lasting 7 hours 47 minutes. As services came back, traffic kept rising. In GitHub’s own words:
“Errors in those services triggered a client-side retry loop that increased traffic during recovery. We had to mitigate that behavior before we could safely restore traffic.”
A retry storm is retries piling onto a system that is already struggling: each layer’s retries multiply, and the extra load arrives exactly when the system is weakest.
What it is
Retrying a failed call is usually sensible: most failures are brief. The trouble starts when several layers of a system each retry. A request passes through a client, a frontend and a backend service before it reaches the database. If each one retries on failure, the retries multiply. It’s also called retry amplification.
The problem
The extra load lands at the worst possible moment. The part that is failing receives the most traffic, which keeps it slow, which causes more failures, which cause more retries.
“Even with a single layer of retries, traffic still significantly increases when errors start.”Marc Brooker, Amazon Builders’ Library
GitHub’s response after its August outage starts with exactly this: “consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions.”
From first principles
Say the database slows down, and the client, the frontend and the service each retry a failed call 3 times. How many attempts reach the database from one click?
- One click: the client sends 1 request to the frontend.
- It fails, so the client retries 3 times: 1 + 3 = 4 attempts reach the frontend.
- Each of those makes the frontend call the service, which also retries 3 times: 4 × 4 = 16 attempts reach the service.
- The service retries its database call 3 times: 16 × 4 = 64 attempts reach the database.
- All 64 hit a database that is already slow, keeping it slow. That loop is the storm.
The general rule:
attempts at the bottom = (retries + 1) ^ layers
| Layers that retry | Retries each | Attempts per click |
|---|---|---|
| 1 | 3 | 4 |
| 2 | 2 | 9 |
| 3 | 2 | 27 |
| 3 | 3 | 64 |
“A single request at the highest layer may produce a number of attempts as large as the product of the number of attempts at each layer.”Google SRE book, “Addressing Cascading Failures” (its own example: 4³ = 64)
The solution
Keep retries, because they let clients ride out brief failures, but stop them multiplying:
- Retry at one layer. AWS: “our best practice is to retry at a single point in the stack.”
- Cap retries with a budget, so retrying stops when most calls are failing.
- Back off, with jitter, so retries don’t arrive as one wave. See backoff with jitter.
- Only retry idempotent calls, or send an idempotency key, so a retry can’t repeat a side effect.
Patterns: real tools already do this
AWS SDKs
[default] retry_mode = standard # default mode max_attempts = 3 # 1 try + 2 retries
Standard mode adds a retry quota: a 500-token bucket where each transient retry costs 14 tokens. When it is empty, the SDK returns the error instead of retrying. Delays use full jitter. (As of October 2026, AWS documents this updated behaviour as opt-in until it becomes the default.)
Envoy
circuit_breakers:
thresholds:
- retry_budget:
budget_percent: { value: 20.0 }
min_retry_concurrency: 3
A retry budget caps concurrent retries as a percentage of active plus pending requests. The values shown are an example.
gRPC
"retryPolicy": { "maxAttempts": 3, ... },
"retryThrottling": { "maxTokens": 10, "tokenRatio": 0.1 }
maxAttempts includes the original call. Each failure costs a token and each success earns tokenRatio; retries pause while tokens are at or below half of maxTokens.
Do
- Retry at one layer, and let the layers below fail fast.
- Use a retry budget or token bucket.
- Use backoff with jitter.
- Use idempotency keys for calls with side effects.
Anti-patterns
- Retrying at every layer. (Google SRE book)
- Retrying without backoff and jitter. (AWS)
- Retrying calls with side effects without an idempotency key. (AWS, Stripe)
- Letting clients keep retrying while the service recovers. (GitHub, Aug 2026)
Trade-off
Fewer retries means some brief blips reach users as errors, and budgets need tuning. AWS also notes that retrying at the highest layer of the stack can waste work from previous calls.
Practise it in two minutes
Predict what happens, watch it move, change the numbers, and get tested again before you forget.
Open Retry storm in the app →Related
Sources
- GitHub, “The August 17 outage, and the work ahead”
- Google SRE book, “Addressing Cascading Failures”
- Marc Brooker, “Timeouts, retries, and backoff with jitter”, Amazon Builders’ Library
- AWS SDKs and Tools Reference Guide, “Retry behavior”
- Envoy, circuit breakers and retry budgets
- gRPC proposal A6, client retries
- Stripe API reference, “Idempotent requests”