Resilience · Pattern
Backoff with jitter
Wait before you retry, wait longer each time, and randomize the wait, so a crowd of clients doesn't hit a recovering service in one wave.
After a failure, don’t retry straight away: wait, wait longer after each failure, and add randomness to every wait so clients spread out instead of arriving together.
What it is
Backoff with jitter is a rule for when to retry a failed request. It has two parts.
1. Backoff: wait, and wait longer each time
Don’t retry immediately. Wait a little after the first failure, longer after the second, longer again after the third. Usually the wait doubles each time, which is called exponential backoff:
attempt 1 fails → wait ~100 ms attempt 2 fails → wait ~200 ms attempt 3 fails → wait ~400 ms
This gives a struggling service room to recover instead of being hammered.
2. Jitter: randomize the wait
Backoff alone has a hidden problem. If 1,000 clients all fail at the same moment, say during a database blip, they all wait exactly 100 ms and retry at the same instant. Then 200 ms, together again. The retries arrive as synchronized waves, which recreate the overload that caused the failure.
Jitter adds randomness, so each client picks a different wait:
without jitter: all retry at exactly 100 ms
→ ▮ one big spike
with jitter: each retries in 0–100 ms
→ ▁▁▁▁▁ a steady trickle
The problem it solves
Failures are rarely isolated. A database blip, a deploy or a network hiccup fails many clients at the same moment, and clients that fail together retry together. If they all wait the same amount, their retries stay in lockstep, wave after wave, aimed at a service that is trying to recover.
“Jitter adds some amount of randomness to the backoff to spread the retries around in time.”Marc Brooker, Amazon Builders’ Library
AWS’s SDK documentation gives a concrete picture: if 1,000 clients all receive a 503 at the same moment, full jitter spreads their first retries evenly across a 50 ms window, instead of all 1,000 retrying at exactly 50 ms.
From first principles
- Why wait at all? Many failures are brief. An immediate retry adds load at the exact moment the service is weakest, and often fails for the same reason.
- Why wait longer each time? If the first retry also failed, the problem is lasting longer. Doubling the wait backs off quickly without making the first retry slow.
- Why randomize? Identical waits keep clients that failed together in sync forever. Randomness is the only thing that breaks the sync.
- Why cap it? Doubling grows fast. A maximum wait stops one request from waiting for minutes.
Put together, this is the formula AWS’s SDKs use, called full jitter:
delay = random(0, 1) × min(cap, base × 2^retry)
In production: how AWS’s SDKs do it
This isn’t theory. The AWS SDKs’ standard retry mode implements exactly this:
| Setting | Value |
|---|---|
| Base delay, transient errors (e.g. HTTP 503) | 50 ms |
| Base delay, throttling errors | 1,000 ms |
| Maximum single delay (cap) | 20 seconds |
| Default max attempts | 3 (1 try + 2 retries) |
For a transient error with 3 attempts, the second attempt waits a random 0–50 ms and the third a random 0–100 ms: about 75 ms added on average. Standard mode also has a retry quota, a token bucket that stops retrying when failures are widespread, so clients fail fast instead of piling on.
As of October 2026, AWS documents this updated behaviour as opt-in (AWS_NEW_RETRIES_2026=true) until it becomes the default; check the guide for your SDK.
And it isn’t only for retries. Brooker’s article recommends adding jitter to timers, periodic jobs and other delayed work too, so scheduled work doesn’t spike at the same second across a fleet.
An everyday picture
A shop’s door jams and 100 people are waiting. Backoff is “step back and try again later, waiting a bit longer each time it fails.” Jitter is “don’t all come back at exactly 9:05; each of you picks a random moment,” so the door isn’t mobbed the instant it reopens.
Do
- Use your SDK or library’s built-in retry mode instead of writing your own.
- Always add jitter to backoff.
- Cap both the delay and the number of attempts.
- Add jitter to scheduled jobs and timers, not just retries.
Don’t
- Retry on a fixed interval (e.g. every second) from many clients.
- Use backoff without jitter: the waves stay synchronized.
- Retry forever.
- Assume backoff alone is enough: retries at several layers still multiply.
How it fits with the rest
Backoff with jitter spreads retries out in time. It doesn’t stop them multiplying when several layers each retry. That’s a retry storm, and the fix there pairs jitter with retrying at a single layer and capping retries with a budget.
Practise it in two minutes
Predict what happens, watch it move, change the numbers, and get tested again before you forget.
Open Backoff with jitter in the app →