Platform Protocols

Resilience · Anti-pattern

Retry storm

When several layers each retry failed calls, one request turns into dozens of attempts, all aimed at the part that is already failing.

By Saravanakumar Raju · 7 min read · Sources verified 4 Oct 2026

Real incident · GitHub · 17 August 2026

GitHub had an outage lasting 7 hours 47 minutes. As services came back, traffic kept rising. In GitHub’s own words:

“Errors in those services triggered a client-side retry loop that increased traffic during recovery. We had to mitigate that behavior before we could safely restore traffic.”
Illustrative shape, not GitHub’s data: errors trigger a retry loop that keeps traffic above capacity until it is mitigated.
© Saravanakumar Raju
In one line

A retry storm is retries piling onto a system that is already struggling: each layer’s retries multiply, and the extra load arrives exactly when the system is weakest.

What it is

Retrying a failed call is usually sensible: most failures are brief. The trouble starts when several layers of a system each retry. A request passes through a client, a frontend and a backend service before it reaches the database. If each one retries on failure, the retries multiply. It’s also called retry amplification.

One click fans out: 1 attempt becomes 4 at the frontend, 16 at the service and 64 at the database.
© Saravanakumar Raju

The problem

The extra load lands at the worst possible moment. The part that is failing receives the most traffic, which keeps it slow, which causes more failures, which cause more retries.

“Even with a single layer of retries, traffic still significantly increases when errors start.”Marc Brooker, Amazon Builders’ Library

GitHub’s response after its August outage starts with exactly this: “consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions.”

From first principles

Say the database slows down, and the client, the frontend and the service each retry a failed call 3 times. How many attempts reach the database from one click?

  1. One click: the client sends 1 request to the frontend.
  2. It fails, so the client retries 3 times: 1 + 3 = 4 attempts reach the frontend.
  3. Each of those makes the frontend call the service, which also retries 3 times: 4 × 4 = 16 attempts reach the service.
  4. The service retries its database call 3 times: 16 × 4 = 64 attempts reach the database.
  5. All 64 hit a database that is already slow, keeping it slow. That loop is the storm.
The same arithmetic, layer by layer: retries multiply, they don’t add.
© Saravanakumar Raju

The general rule:

attempts at the bottom = (retries + 1) ^ layers
Layers that retryRetries eachAttempts per click
134
229
3227
3364
“A single request at the highest layer may produce a number of attempts as large as the product of the number of attempts at each layer.”Google SRE book, “Addressing Cascading Failures” (its own example: 4³ = 64)

The solution

Keep retries, because they let clients ride out brief failures, but stop them multiplying:

  1. Retry at one layer. AWS: “our best practice is to retry at a single point in the stack.”
  2. Cap retries with a budget, so retrying stops when most calls are failing.
  3. Back off, with jitter, so retries don’t arrive as one wave. See backoff with jitter.
  4. Only retry idempotent calls, or send an idempotency key, so a retry can’t repeat a side effect.

Patterns: real tools already do this

AWS SDKs

[default]
retry_mode   = standard   # default mode
max_attempts = 3          # 1 try + 2 retries

Standard mode adds a retry quota: a 500-token bucket where each transient retry costs 14 tokens. When it is empty, the SDK returns the error instead of retrying. Delays use full jitter. (As of October 2026, AWS documents this updated behaviour as opt-in until it becomes the default.)

Envoy

circuit_breakers:
  thresholds:
  - retry_budget:
      budget_percent: { value: 20.0 }
      min_retry_concurrency: 3

A retry budget caps concurrent retries as a percentage of active plus pending requests. The values shown are an example.

gRPC

"retryPolicy":    { "maxAttempts": 3, ... },
"retryThrottling": { "maxTokens": 10, "tokenRatio": 0.1 }

maxAttempts includes the original call. Each failure costs a token and each success earns tokenRatio; retries pause while tokens are at or below half of maxTokens.

Do

  • Retry at one layer, and let the layers below fail fast.
  • Use a retry budget or token bucket.
  • Use backoff with jitter.
  • Use idempotency keys for calls with side effects.

Anti-patterns

  • Retrying at every layer. (Google SRE book)
  • Retrying without backoff and jitter. (AWS)
  • Retrying calls with side effects without an idempotency key. (AWS, Stripe)
  • Letting clients keep retrying while the service recovers. (GitHub, Aug 2026)

Trade-off

Fewer retries means some brief blips reach users as errors, and budgets need tuning. AWS also notes that retrying at the highest layer of the stack can waste work from previous calls.

Practise it in two minutes

Predict what happens, watch it move, change the numbers, and get tested again before you forget.

Open Retry storm in the app →

Related

Sources

  1. GitHub, “The August 17 outage, and the work ahead”
  2. Google SRE book, “Addressing Cascading Failures”
  3. Marc Brooker, “Timeouts, retries, and backoff with jitter”, Amazon Builders’ Library
  4. AWS SDKs and Tools Reference Guide, “Retry behavior”
  5. Envoy, circuit breakers and retry budgets
  6. gRPC proposal A6, client retries
  7. Stripe API reference, “Idempotent requests”