Retrying is instinctive. Something didn’t work, so you try again. We do it when the card terminal beeps at us, when a door sticks, when a page hangs halfway through loading. Code is no different, and I can’t think of a codebase I’ve worked on that didn’t retry something somewhere.
Sometimes it’s as naive as this:
Response callWithRetry(Request request) throws Exception {
Exception last = null;
for (int attempt = 1; attempt <= 3; attempt++) {
try {
return client.send(request);
} catch (Exception e) {
last = e;
Thread.sleep(1000);
}
}
throw last;
}
Sometimes it’s a Resilience4j Retry with exponential backoff and jitter, wrapped around a circuit breaker and a time
limiter, configured per dependency in YAML, with its own dashboard. Most of the time it’s something in between: a
@Retryable(maxAttempts = 3) on a method, an SDK that retries without asking, or a line in the service mesh config
that nobody on the team remembers adding. Underneath, they all do the same thing. The call fails, you wait a little,
and you try again.
And most of the time that works fine. A packet gets dropped, a pod gets rescheduled, a pooled connection turns out to be dead, and the second attempt goes through without the user ever noticing. If you only look at one client making one call, it’s hard to see a downside.
The problem is that nobody runs just one client. The retry policy runs in every instance of every caller, it kicks in on failure, and failures tend to arrive together. When the database has a bad thirty seconds, it has them for everyone, so every retry loop wakes up at roughly the same moment. Each of those loops was written to give one request a better chance. Put them all together and the struggling dependency gets two or three times its usual traffic, at the exact moment it can least handle it.
So the question I care about is what happens when every caller retries at once. That’s what this article is about, and it breaks down into these cases:
- several layers each retry, and the attempts multiply
- a
400gets retried three times and fails anyway, only slower - a timeout fires after the server already did the work
- thousands of clients back off on the same schedule and return together
- the downstream is overloaded, and retries add load
- the retry library does something different from what you assumed
- the server sends
Retry-Afterand every client obeys it to the second - a
POSTgets retried with a fresh idempotency key on every attempt - a queue consumer keeps retrying a message that will never succeed
- the original problem goes away, and the retries keep the system down anyway
Each one has its own section below, with what I’d do about it and, where a library can do the work, how that looks in Resilience4j. The tests and the checklist at the end tie it all together.
Handling the flaky-packet case is the easy part. Most of what follows is about the others.
Backoff is older than your codebase
None of this is new. I think the history is worth a couple of minutes, though, because it explains why the fixes look the way they do.
One of the first well-known retry problems came from a radio network. ALOHAnet, built at the University of Hawaii in the early 1970s, let stations transmit whenever they wanted. If two transmissions overlapped, both were lost and both stations had to send again, and if they each waited the same fixed amount of time, they would simply collide again. The answer was to have each station wait a random amount of time, so randomness was part of the solution from day one.
Ethernet built on that. In the 1976 paper that describes it, Metcalfe and Boggs explain that after each collision the controller waits a random interval whose average doubles every time, which they call Binary Exponential Backoff.[15] There’s a line in that paper I really like. They point out that a station could “usurp the Ether” by not adjusting its retransmission interval as traffic grows, and then add that “both practices are now prohibited by low-level software in each station.” Backing off was simply part of the deal when you shared a wire with other people, and a station that refused to do it was considered broken.
TCP had to learn the same thing the hard way. In October 1986 the early internet went through its first congestion collapse, and throughput between Lawrence Berkeley Laboratory and UC Berkeley, two sites 400 yards apart, fell from 32 Kbps to 40 bps.[16] Senders were retransmitting lost packets into a network that was losing them precisely because it was overloaded. Among the fixes Van Jacobson and Mike Karels published in 1988 was exponential backoff for the retransmit timer. If you squint a bit, that collapse looks a lot like a retry storm.
For a long time, application developers didn’t need to care about any of this. TCP handled retransmission underneath them, and plenty of applications were a single process talking to a database in the same data center. Then, through the 2000s, systems got split into services. A single user action started crossing several network hops, any of which could fail on its own, and each hop grew its own retry loop. Those loops were mostly written from scratch, with a fixed delay and no randomness, by people who had never had a reason to read the Ethernet paper.
The lessons came back slowly, and mostly through outages. Michael Nygard’s Release It! (2007) collected the patterns teams had worked out the hard way and gave them names that stuck: timeouts, circuit breakers, bulkheads.[17] Netflix later packaged several of them into Hystrix, which is how a lot of JVM developers first came across them. Hystrix has been in maintenance mode since 2018, and its README now sends new projects to Resilience4j.[18] Jitter and retry budgets took longer. They only became standard advice in the mid-2010s, when Marc Brooker’s post on jitter[1] and Google’s SRE book[4] explained them to backend engineers who were running into the problem TCP had hit thirty years earlier.
What I find interesting is that at the lower layers, backoff and randomness arrived together as one idea, along with the understanding that you back off because other people share the network with you. Application-level retries mostly kept the loop and lost the rest. A lot of this article is just those older lessons, moved up the stack.
Count the attempts at the bottom
I’ll start with some arithmetic, since it’s the step people usually skip.
Say a mobile app makes up to 3 attempts. It calls an API gateway, which also makes up to 3. The gateway calls an orders service, and the orders service’s HTTP client makes up to 3 attempts to payments. Payments talks to the database through a driver that retries 3 times as well.
mobile app 3 attempts
API gateway 3 attempts per request it receives
orders service 3 attempts per request it receives
payments service 3 attempts per request it receives
database up to 3 × 3 × 3 × 3 = 81 attempts per tap
Nobody ever decided on 81. Each team picked 3, which feels reasonable, and the 81 only shows up when you multiply decisions made in four different repositories.
When everything is healthy this costs nothing, because the first attempt succeeds and nothing gets retried. The multiplier only kicks in when the bottom layer is failing, which is the worst possible time for it. Marc Brooker describes the single-layer case well: with N retries and a downstream that fails every call, the system does 1 + N times the work.[2] Add more layers and it compounds.
Google’s SRE book calls this a combinatorial explosion, and its rule is one I agree with: only the layer directly above the one rejecting requests should retry.[4] I’ll come back to that. For now, try a simple exercise. Draw the call path from the user down to the deepest dependency and write the retry setting next to each hop, including the ones you didn’t write yourself, like SDK defaults, HTTP client defaults, database drivers, the service mesh, the load balancer and the message broker. Then multiply them.
If there’s a hop where nobody knows the setting, you’ve already learned something useful.
Some failures should never be retried
The loop at the top of this article catches Exception. So it will happily retry a 400 Bad Request three times,
sleeping a second in between, and then fail exactly the way it was always going to fail, just slower.
Before a client decides how to retry, it needs to know what actually failed. I find it helps to put failures into three groups:
- The request was never processed. Connection refused, DNS failure, connect timeout, an HTTP/2
REFUSED_STREAM. The server never saw the request, so retrying is safe for any method. - The request was processed and rejected. Validation errors, auth failures, not found, business rules. Retrying just gets you the same answer again.
- Nobody knows. The read timed out after the request went out, the connection was reset halfway through the
response, or a gateway returned
504. The server might have done nothing, some of the work, or all of it.
Most of the damage comes from the third group, and it’s bigger than it looks.
Here’s roughly how I’d treat the common HTTP cases:
| Failure | Retry? | Notes |
|---|---|---|
| connection refused, DNS failure, connect timeout | yes | the request never reached the application |
| read timeout after sending | only if the operation is idempotent | the outcome is unknown |
400, 422 | no | the request itself is wrong |
401 | once, after refreshing credentials | then stop |
403, 404 | no | unless the API documents eventual consistency for that resource |
409 | depends | an in-progress idempotent request with Retry-After can be retried; a business conflict can’t |
429 | yes, after waiting | honor Retry-After; the server is asking you to slow down |
500 | cautiously | often a bug that will fail again; never blindly for non-idempotent calls |
502, 503 | yes, with backoff | the usual transient or overload signals |
504 | only if the operation is idempotent | the gateway gave up; the upstream may have finished |
People are often surprised by the 504 row. A gateway timeout only tells you that the gateway stopped waiting. It
tells you nothing about whether the service behind it stopped working, and the order may well have been committed two
seconds after the gateway gave up.
HTTP is quite careful about this. RFC 9110 says a client SHOULD NOT automatically retry a request with a non-idempotent
method unless it knows the request semantics are actually idempotent, or has some way to detect that the original
request was never applied.[9] HTTP/2 goes further and gives servers two ways to say “I didn’t process this
one”: the last stream ID in a GOAWAY frame, and the REFUSED_STREAM error code. Requests covered by either can be
retried even if they’re POSTs.[10] That’s group 1 at the transport level, and it’s about the only place you
get that guarantee without doing any work yourself.
gRPC makes the same distinction. Without a configured retry policy, it only retries transparently when it’s sure the server application never saw the call, and anything beyond that needs an explicit policy that lists the retryable status codes.[6]
There’s arguably a fourth group, where the server makes the call: “I’m overloaded, don’t retry.” Google’s SRE book
describes backends that return a separate error when they can see retries piling up, so callers back off instead of
adding to the pile.[4] Most HTTP APIs have nothing like this, but a stable error code in the body of a 503 is
enough to build it.
In Resilience4j, retryOnException and retryOnResult are where this classification lives, and I’d always set both.
If you leave out the exception predicate, the default is throwable -> true, which means every exception gets retried,
including whatever your client throws for a 400.[20] With the JDK HttpClient I’d end up with two
configurations, one for calls that are safe to repeat and one for calls that aren’t (all the examples here use
Resilience4j 2.x):
RetryConfig unsafeCallRetry = RetryConfig.<HttpResponse<String>>custom()
.maxAttempts(3)
.retryOnException(e -> e instanceof ConnectException
|| e instanceof HttpConnectTimeoutException)
.retryOnResult(response -> response.statusCode() == 429)
.build();
RetryConfig idempotentCallRetry = RetryConfig.<HttpResponse<String>>custom()
.maxAttempts(3)
.retryOnException(e -> e instanceof IOException)
.retryOnResult(response -> Set.of(429, 502, 503, 504).contains(response.statusCode()))
.failAfterMaxAttempts(true)
.build();
The first one only retries the “never processed” group, plus 429, which normally means the server turned the request
away before doing anything with it. The second one also retries read timeouts (an HttpTimeoutException is an
IOException) and 504, since repeating the call is harmless. failAfterMaxAttempts(true) tells Resilience4j to throw
MaxRetriesExceededException if the last attempt still comes back with a retryable status. Without it, that final
503 is returned like any other response, and code that doesn’t check the status will happily treat it as one.
A timeout tells you nothing about the outcome
This is where retries and idempotency run into each other, and it’s where most duplicate side effects come from.
A client sends POST /payments. The server gets it, validates it, calls the payment provider, writes the row, and
starts serializing the response. The client’s read timeout fires at 5 seconds, and the response would have arrived at
5.2.
As far as the client knows, the request failed. As far as the server knows, it succeeded. So the retry loop sends it again.
Without an idempotency key, you now have two payments. With one, it comes down to how well the server handles the second request. That’s a long topic, and it’s the reason I wrote Idempotency Is Easy Until the Second Request Is Different. Briefly, the server needs an atomic way to decide who owns execution, a way to recognize a replay of the same command, and a plan for a retry that shows up while the first attempt is still running.
That last one is what retries cause most often. A client with a 2-second timeout calling an endpoint that sometimes takes 3 seconds will regularly send a retry while the original is still in flight, and the server ends up running the same operation twice at the same time. If its idempotency check is “look up the key, then insert” instead of an atomic insert, both copies go ahead. The section on who owns execution in that article explains why.
It also changes how you should pick timeouts. A client timeout should come from how long the operation can legitimately take, based on real latency numbers, plus some headroom. Set it near the median and you guarantee that a big chunk of requests that would have succeeded get retried anyway. People usually defend these timeouts as failing fast, but in practice what they mostly do is create duplicate work.
And the server has to assume it’ll keep working after the client has given up. Most frameworks don’t cancel a handler when the client disconnects, and even the ones that do can’t take back a call that already went out to a payment provider. Any retry after a timeout might be a second copy of the operation running alongside the first, and the server needs to be built with that in mind.
Timeouts have to nest
Retries multiply the number of attempts. Timeouts decide whether anyone is still around to see the result.
Here’s a configuration I suspect exists, in some form, in a lot of systems:
mobile app → API gateway timeout 10s, 3 attempts
API gateway → orders service timeout 30s
orders service → payments service timeout 5s, 3 attempts, backoff up to 2s
payments service → provider timeout 20s
Now the provider slows down and starts taking 25 seconds to respond.
Payments will wait up to 20 seconds for each provider call. Orders gives up on payments after 5 seconds and retries, then gives up and retries again, while each abandoned request is still sitting in payments waiting on the provider. After 10 seconds the mobile app gives up on the gateway and retries, which kicks off a whole new chain. The gateway’s 30-second timeout never comes into play, because by then nobody is listening.
So a single tap can lead to as many as nine provider calls in flight for the same payment within about half a minute. None of those responses will ever reach the person who tapped, and every one of them has an unknown outcome.
The rules I try to stick to are simple enough:
- Inner timeouts are shorter than outer ones.
- A layer’s total retry time (attempts × per-attempt timeout + backoff) has to fit inside its caller’s timeout, otherwise the later attempts are work nobody is waiting for.
- Even better, pass the deadline down the chain instead of configuring every hop separately.
With deadline propagation, the caller says “I need an answer by 10:00:03.500,” and each hop looks at how much time is left before deciding what to do. gRPC supports this out of the box: the client sets a deadline, servers can check whether it has already passed, and a server that calls other services can pass the original deadline along.[7] Over plain HTTP you have to carry it yourself, usually in a header, either as an absolute timestamp or as the remaining budget in milliseconds.
Once there’s a deadline, the retry decision becomes a lot more sensible:
for (int attempt = 1; attempt <= maxAttempts; attempt++) {
Duration remaining = Duration.between(clock.instant(), deadline);
if (remaining.compareTo(minUsefulAttempt) < 0) {
break;
}
Duration attemptTimeout = min(perAttemptTimeout, remaining);
// send with attemptTimeout, classify the failure, back off
}
If there are 80 milliseconds left and the call normally takes 200, there’s no point sending it. Failing straight away is cheaper for you and for the dependency.
Backoff buys time
The fixed Thread.sleep(1000) in the opening loop has a very specific problem. Every client that failed at the same
moment retries at the same moment, one second later, and then again a second after that. If the dependency couldn’t
cope with the first wave, an identical second wave isn’t going to go any better.
Exponential backoff makes the delay grow with each attempt, say 100ms, 200ms, 400ms, 800ms, up to some cap. A dependency that’s briefly overloaded gets some room to recover, and the retries get spread over a longer window.
It helps less than most people think, though, and Brooker’s “What is Backoff For?” explains why better than I could.[3] Backoff postpones work. It only reduces the total amount of work when the clients are a small, fixed group, each sending one request after another, like a fleet of workers polling a queue. A worker that’s sleeping isn’t sending its next request either, so the system actually gets a break.
Most public APIs don’t look like that. They have a huge, open-ended population of callers, each deciding on its own when to send its first request. A page that a million people load once a day doesn’t get less traffic just because some clients are backing off, since new visitors keep showing up at the same rate. In that situation backoff smooths out short spikes, which is useful, but during a long overload it only pushes the retries later, where they pile up behind the next wave of first attempts.
So you need backoff, and it deals with short blips well, but it does nothing to limit how much extra load retries add during a sustained outage. For that you need a budget, which I’ll get to in a couple of sections.
Jitter breaks up the crowd
Exponential backoff without any randomness still keeps clients in lockstep. Ten thousand clients that failed at t=0
will all retry at t=100ms, then at t=300ms, then at t=700ms. The load arrives as spikes with quiet gaps in
between, and it’s the spikes that do the damage.
Jitter adds randomness to each delay so the clients drift apart. My default is what Brooker’s 2015 AWS post calls “full jitter”, which just picks a random delay between zero and the capped exponential value.[1]
long backoffMillis(int retry) {
long exponential = BASE_MILLIS << Math.min(retry - 1, 16);
long capped = Math.min(MAX_BACKOFF_MILLIS, exponential);
return ThreadLocalRandom.current().nextLong(capped + 1);
}
In his simulations, both full jitter and “decorrelated jitter” cut the total client work substantially compared to no jitter at all. Full jitter did slightly less work, and “equal jitter” (keep half the delay, randomize the other half) came out worst of the three.[1] gRPC’s built-in retries use a narrower jitter of plus or minus 20% around the backoff value.[6] Any of them is far better than nothing, and honestly the exact formula matters a lot less than having some randomness in there at all.
Resilience4j’s default Retry has neither backoff nor jitter; it waits a flat 500ms between
attempts.[20] You get both from an IntervalFunction:
IntervalFunction backoff = IntervalFunction.ofExponentialRandomBackoff(
Duration.ofMillis(100), 2.0, 0.5, Duration.ofSeconds(5));
RetryConfig config = RetryConfig.custom()
.maxAttempts(3)
.intervalFunction(backoff)
.build();
Or, with the Spring Boot starter:
resilience4j.retry:
instances:
payments:
maxAttempts: 3
waitDuration: 100ms
enableExponentialBackoff: true
exponentialBackoffMultiplier: 2
exponentialMaxWaitDuration: 5s
enableRandomizedWait: true
randomizedWaitFactor: 0.5
There are two details here that I only found by reading the source.[29] First, the randomization factor spreads the delay evenly around the exponential value, so with 0.5 a base delay of 800ms can come out anywhere between 400ms and 1200ms. That’s the plus-or-minus kind of jitter, closer to what gRPC does than to full jitter. Second, the cap is applied after the randomization. Once the exponential value goes past the cap, most of the randomized delays land above it and get clipped to exactly the cap, which puts clients back in step. With a 1-second initial interval, a 5-second cap and a factor of 0.5, the fourth retry has a base of 8 seconds, draws something between 4 and 12, and ends up at exactly 5 seconds about seven times out of eight.
You can either keep the number of attempts low enough that you never reach the cap, or write your own
IntervalFunction that caps first and randomizes second, which gets you back to full jitter:
IntervalFunction fullJitter = attempt -> {
long exponential = 100L << Math.min(attempt - 1, 16);
return ThreadLocalRandom.current().nextLong(Math.min(5_000L, exponential) + 1);
};
The same lockstep problem turns up in places that have nothing to do with retry loops:
- mobile apps that all reconnect the moment connectivity comes back after an outage
- cron jobs scheduled at
0 * * * *for hundreds of tenants - cache entries written together during a deploy, with the same TTL, that then expire together
- clients that refresh auth tokens on a fixed interval after startup, all of them started by the same rollout
The fix is the same in every case: add a random offset so the fleet stops moving in step.
Budgets cap the multiplier
A maximum number of attempts per request is a limit you definitely need, and it’s usually the only one retry code has. It caps what one request can cost, but it says nothing about what the client as a whole costs the dependency when everything is failing. At a 100% failure rate, “3 attempts” means three times the load, which is the opposite of what you want at that moment.
A retry budget limits retries as a fraction of total traffic. While failures are rare, every failure can be retried. Once they become common, retries dry up, and the client goes back to sending roughly its normal number of first attempts.
The token bucket version is short:
final class RetryBudget {
private final double maxTokens;
private final double tokensPerSuccess;
private double tokens;
RetryBudget(double maxTokens, double tokensPerSuccess) {
this.maxTokens = maxTokens;
this.tokensPerSuccess = tokensPerSuccess;
this.tokens = maxTokens;
}
synchronized void recordSuccess() {
tokens = Math.min(maxTokens, tokens + tokensPerSuccess);
}
synchronized boolean hasTokens() {
return tokens >= 1;
}
synchronized boolean tryAcquireRetry() {
if (tokens < 1) {
return false;
}
tokens -= 1;
return true;
}
}
With tokensPerSuccess = 0.1, each success adds a tenth of a token and each retry costs a full one, which over time
caps retries at about 10% of successful calls. When the downstream falls over, successes stop coming in, the bucket
empties after a handful of retries, and the client sends only first attempts until things recover.
You’ll find the same idea in a lot of places:
- gRPC’s
retryThrottlingconfig keeps a token count per server. Failures take a token away, successes addtokenRatio, and retries stop while the count is below half ofmaxTokens.[6] - The AWS SDKs’ standard retry mode has a retry quota that works the same way. It never delays the first request, only retries.[8]
- Google’s SRE book describes a per-client budget where a request is retried only while retries are under 10% of that client’s traffic, and says this brings the growth from retries down from 3x to about 1.1x in the general case.[4]
- Finagle clients retry out of a budget that allows roughly 20% of requests to be retried, plus 10 retries per second so that new or quiet clients can still retry.[19]
- Envoy can cap concurrent retries at a percentage of active requests, 20% by default, once you configure a retry budget on the cluster.[28]
Resilience4j doesn’t come with a retry budget, but its predicates and events make it easy enough to add the one above. You check the budget when deciding whether a failure is retryable, spend a token when a retry actually happens, and record successes around the call:
RetryBudget budget = new RetryBudget(10, 0.1);
RetryConfig config = RetryConfig.<HttpResponse<String>>custom()
.maxAttempts(3)
.intervalFunction(fullJitter)
.retryOnException(e -> e instanceof IOException && budget.hasTokens())
.retryOnResult(r -> RETRYABLE_STATUSES.contains(r.statusCode()) && budget.hasTokens())
.build();
Retry paymentsRetry = Retry.of("payments", config);
paymentsRetry.getEventPublisher().onRetry(event -> budget.tryAcquireRetry());
HttpResponse<String> response = paymentsRetry.executeCallable(() -> {
HttpResponse<String> r = http.send(request, BodyHandlers.ofString());
if (!RETRYABLE_STATUSES.contains(r.statusCode())) {
budget.recordSuccess();
}
return r;
});
The token is spent in the onRetry event because that event only fires when a retry is actually about to happen.
Resilience4j also evaluates the predicates after the last attempt, when no retry follows, so spending tokens there would
overcharge. The Retry instance, and the budget with it, should be shared by everything that calls the same
dependency. If you create one per request, every request starts with a full bucket and the budget doesn’t limit
anything.
There’s one caveat from Brooker’s simulations: each client estimates its own budget, and small clients make noisy estimates.[2] A thousand short-lived serverless functions, each with its own full bucket, end up behaving a lot like a thousand clients with plain “N retries,” because none of them sees enough failures to drain its bucket. If your callers are many and short-lived, a budget in a shared proxy or sidecar sees more traffic and makes better decisions than one inside each function.
Pick one layer to retry
Back to those 81 attempts. The change that pays off most is usually picking one layer to own retries for each hop and switching them off everywhere else.
That owner should normally be the layer right above the failing dependency, because it has the most context. It knows which errors are transient for that dependency, whether the operation is idempotent, and how much of its deadline is left. Layers further up only see “orders returned 503” and have no idea three retries already happened underneath.
When the owning layer gives up, it should tell its callers to stop as well. That could be a 503 with an error code
like DEPENDENCY_UNAVAILABLE that’s documented as non-retryable, or a degraded response if there’s a sensible one.
Without some signal like that, the next layer up treats the failure as transient and starts a loop of its own.
The retries you didn’t write are usually hiding in places like these:
- HTTP client libraries. Some retry certain failures on their own, usually connection-level ones. Check the defaults for the one you actually use, because they differ between libraries and sometimes between versions of the same library.
- Cloud SDKs. The AWS SDKs’ standard mode defaults to 3 attempts in total.[8]
- Service meshes. Mesh-level retries are one YAML block away, and easy to forget about while the application retries too.
- Load balancers and reverse proxies. nginx’s
proxy_next_upstreampasses a failed request on to the next upstream server on errors and timeouts. Since 1.9.13 it won’t do that forPOST,LOCKorPATCHonce the request has been sent upstream, unless you add thenon_idempotentoption.[12] It’s worth grepping your configs for that one. - Database drivers and ORMs. Some reconnect and re-run statements after a failover. Find out whether yours does, and for which statements.
- Framework annotations, like a
@Retryablemethod that calls another@Retryablemethod. - Message brokers. Redelivery is a retry policy as well, more on that below.
- People. Someone staring at a spinner is going to press the button again. If the UI doesn’t disable the submit button, or doesn’t reuse the same idempotency key on the second click, that user is one more retry layer, with no backoff whatsoever.
Read the defaults before trusting a library
I would use a library for all of this. The hand-written snippets in this article are there to show how things work, and a maintained library will handle edge cases that a loop written on a Friday afternoon won’t. Still, a library’s retry is a retry policy like any other, and its defaults are choices someone else made for the general case. They vary a lot more than I expected:
| Library | Attempts with default settings | Default delay | Retries by default |
|---|---|---|---|
Resilience4j Retry | 3 in total | fixed 500ms | every exception |
Spring Retry @Retryable (archived) | 3 in total | fixed 1s | every exception |
Spring Framework 7 @Retryable | 4 (1 + maxRetries = 3) | fixed 1s, jitter off | every exception |
| Polly v8 retry strategy | 4 (1 + MaxRetryAttempts = 3) | constant 2s, jitter off | every exception except OperationCanceledException |
.NET AddStandardResilienceHandler() | 4 (1 + 3 retries) | exponential from 2s, jitter on | 5xx, 408, 429, HttpRequestException and timeouts, for every HTTP method |
tenacity @retry (Python) | unlimited | none | every exception |
| go-retryablehttp | 5 (1 + RetryMax = 4) | exponential from 1s to 30s, no jitter | connection errors, 429, and 5xx except 501, for every HTTP method |
Sources for each row are in the references.[20][22][23][24][25][26][27]
The first thing I notice in that table is that “3” means three attempts in some libraries and four in others. Even
Spring’s two @Retryable annotations disagree: Spring Retry’s maxAttempts includes the first call, and Spring
Framework 7’s maxRetries doesn’t. Spring Retry is now archived in favor of the Spring Framework 7 annotation, so
plenty of teams will migrate from one to the other and go from three attempts to four without noticing.
Most of them also retry every exception unless you say otherwise. That includes the exception for a 400, a failed
deserialization, and a NullPointerException in your own code.
Jitter is usually something you have to turn on. The .NET standard handler is the odd one out here, with exponential
backoff and jitter, a circuit breaker, and both per-attempt and total timeouts out of the box. It also retries POST by
default, until you call DisableForUnsafeHttpMethods().[25] go-retryablehttp retries regardless of method as
well, and when it gets a 429 or 503 with Retry-After, its default backoff uses the server’s value exactly as sent,
with no jitter.[27] That’s precisely the synchronized return I describe in the next section.
To be fair, tenacity’s wait_random_exponential implements full jitter, and its docs link to the AWS post that named
it.[26] Most of these libraries can do the right thing; they just don’t all do it by default.
Order matters in Resilience4j
Resilience4j splits resilience into separate pieces (Retry, CircuitBreaker, TimeLimiter, Bulkhead and
RateLimiter), and how you nest them changes what happens during an outage. With the Spring Boot annotations the
default order is Retry ( CircuitBreaker ( RateLimiter ( TimeLimiter ( Bulkhead ( Function ) ) ) ) ), so every attempt
goes through the circuit breaker and gets its own time limit.[21] With the functional style you choose the
order yourself, and each with... call wraps whatever came before it:
Supplier<HttpResponse<String>> call = Decorators.ofSupplier(() -> send(request))
.withCircuitBreaker(paymentsBreaker)
.withRetry(paymentsRetry)
.decorate();
(send here wraps the checked IOException in an UncheckedIOException, because a Supplier can’t throw checked
exceptions.)
With the retry outside the breaker, once the breaker opens every attempt fails immediately with
CallNotPermittedException. The default exception predicate considers that retryable, so Retry waits, tries again,
hits the same open breaker, and keeps doing that until it runs out of attempts. Adding the exception to
ignoreExceptions, which applies on top of any predicate you’ve configured, makes an open breaker end the call right
away:
RetryConfig config = RetryConfig.<HttpResponse<String>>custom()
.maxAttempts(3)
.intervalFunction(fullJitter)
.retryOnException(e -> e instanceof IOException || e instanceof UncheckedIOException)
.ignoreExceptions(CallNotPermittedException.class)
.build();
A TimeLimiter inside the Retry is a per-attempt timeout. It doesn’t limit the whole call including backoff, so you
still need an overall deadline, either checked in the retry predicates or enforced by the caller. And when it fires,
cancelRunningFuture only cancels the local future. The remote server carries on working, as I described in the
timeout section.
Retry-After deserves respect, and jitter
When a server sends 429 Too Many Requests or 503 Service Unavailable with a Retry-After header, it’s telling the
client how long to wait.[9][11] Clients should listen. A lot of them don’t, especially hand-written
ones, and a client that responds to a 429 by trying again 100ms later is the reason the server needed rate limiting
in the first place.
There are two details people tend to miss. The first is that Retry-After can be either a number of seconds or an HTTP
date:[9]
Retry-After: 30
Retry-After: Fri, 25 Sep 2026 10:15:30 GMT
A client that only parses integers will see the date form as garbage, so you need to decide what happens then. Falling
back to normal backoff is fine; retrying immediately isn’t. You should clamp the value too. A server bug that sends
Retry-After: 86400 shouldn’t park your worker for a whole day, and if the value is longer than the time you have
left, you might as well fail now.
Optional<Duration> retryAfter(HttpResponse<?> response, Instant now) {
return response.headers().firstValue("Retry-After")
.map(String::trim)
.flatMap(value -> parseSeconds(value).or(() -> parseHttpDate(value, now)))
.map(delay -> delay.isNegative() ? Duration.ZERO : delay)
.map(delay -> delay.compareTo(MAX_RETRY_AFTER) > 0 ? MAX_RETRY_AFTER : delay);
}
Optional<Duration> parseSeconds(String value) {
try {
return Optional.of(Duration.ofSeconds(Long.parseLong(value)));
} catch (NumberFormatException e) {
return Optional.empty();
}
}
Optional<Duration> parseHttpDate(String value, Instant now) {
try {
var at = ZonedDateTime.parse(value, DateTimeFormatter.RFC_1123_DATE_TIME);
return Optional.of(Duration.between(now, at.toInstant()));
} catch (DateTimeParseException e) {
return Optional.empty();
}
}
The second detail is my favorite. If a server tells ten thousand clients Retry-After: 30 during an incident, and every
one of them waits exactly thirty seconds, the server has just scheduled its own thundering herd for thirty seconds
later. Clients should add a bit of jitter on top of Retry-After, and servers can help by randomizing the value they
send, say anywhere from 20 to 40 seconds instead of a flat 30.
In Resilience4j, intervalBiFunction gets the attempt number and either the exception or the result, so it can read
Retry-After from the response and fall back to normal backoff for everything else. It replaces intervalFunction,
and configuring both throws an IllegalStateException.[20]
IntervalBiFunction<HttpResponse<String>> retryAfterOrBackoff = (attempt, outcome) -> {
long backoff = fullJitter.apply(attempt);
if (outcome.isLeft()) {
return backoff;
}
return retryAfter(outcome.get(), Instant.now())
.map(delay -> delay.toMillis() + ThreadLocalRandom.current().nextLong(1_000))
.orElse(backoff);
};
RetryConfig config = RetryConfig.<HttpResponse<String>>custom()
.maxAttempts(3)
.retryOnResult(r -> (r.statusCode() == 429 || r.statusCode() == 503) && retryAfterFitsDeadline(r))
.intervalBiFunction(retryAfterOrBackoff)
.build();
Note that the deadline check goes in the predicate. By the time the interval function runs, Resilience4j has already decided to retry, and it’s too late to say “don’t bother, this would take longer than the time I have left.”
In the idempotency article I suggested 409 Conflict with Retry-After for a request that arrives while an earlier
attempt with the same key is still running. The same thinking applies there: telling the client when to come back is
better than letting it spin, and that value should be jittered too.
The idempotency key belongs outside the loop
To retry a POST safely you need an idempotency key, and that key has to stay the same across every attempt of the
same operation. It sounds obvious, and yet this bug still makes it to production:
for (int attempt = 1; attempt <= maxAttempts; attempt++) {
var request = HttpRequest.newBuilder(paymentsUri)
.header("Idempotency-Key", UUID.randomUUID().toString())
.POST(body)
.build();
// send, classify, back off
}
With a new key on every attempt, each retry looks like a brand-new payment. The server’s idempotency layer can be perfect and it still won’t help, because the client has told it these are different operations.
The key should be created once, at the point where the client decides to perform the operation:
String idempotencyKey = paymentAttempt.idempotencyKey();
for (int attempt = 1; attempt <= maxAttempts; attempt++) {
var request = HttpRequest.newBuilder(paymentsUri)
.header("Idempotency-Key", idempotencyKey)
.POST(body)
.build();
// send, classify, back off
}
Reading the key from paymentAttempt is intentional. If the key only lives in memory, a client that crashes after the
first attempt comes back without it, and whatever picks up the work after the restart (a scheduled job, a queue
redelivery) will generate a new one. For operations that matter, store the key with the record that represents the
intent, like the order or the payment attempt, before the first call goes out. That way every retry, from any process,
uses the same key.
The body has to stay the same as well. If a retry re-reads the cart and the total has changed in the meantime, the second request carries the same key with a different amount. A well-built server will reject that, as I described in same key, different command. Build the request once and send the same thing every time.
Libraries make this easier to get wrong, because the retry boundary is harder to see. With Resilience4j, everything inside the decorated lambda runs again on every attempt, so the request and its key have to be built outside it:
HttpRequest request = HttpRequest.newBuilder(paymentsUri)
.header("Idempotency-Key", paymentAttempt.idempotencyKey())
.POST(BodyPublishers.ofString(body))
.build();
HttpResponse<String> response = paymentsRetry.executeCallable(
() -> http.send(request, BodyHandlers.ofString()));
Move UUID.randomUUID() or a fresh read of the cart into that lambda and you’ve got both bugs back. Annotation-based
retries have the same trap, just somewhere less obvious: @Retryable re-runs the entire annotated method, so a key
generated inside that method is a different key on every attempt. Generate it in the caller or, better, load it from
the stored payment attempt.
The same applies one level down, to the calls your service makes on the user’s behalf. When you retry a call to a payment provider, the provider’s idempotency key should be derived from your own operation ID, so it’s identical on every attempt.
Queues retry too
HTTP retries are visible in the code. Queue retries often aren’t, because the broker does them for you.
When a consumer fails to process a message, most setups make the message available again. An SQS message becomes visible again after its visibility timeout, a Kafka consumer that doesn’t commit its offset sees the record again, and a RabbitMQ message that’s nacked with requeue goes back on the queue. Each of these is a retry policy with the same problems as the HTTP kind, usually with less oversight.
Two patterns show up over and over.
Poison messages. Some messages fail every time because of what’s in them: a missing field, an unknown enum value, a reference to something that’s been deleted. Immediate redelivery turns that into a tight loop, and on a Kafka partition that’s processed in order it also blocks everything behind it. The fix isn’t exciting, but you need it: cap the number of attempts, then move the message somewhere a person can look at it, usually a dead-letter queue. In the enums article I mentioned a webhook handler that returns 500 on an unexpected enum value. If the sender retries webhooks, that’s the same poison-message loop, only across a network boundary.
Downstream outages. Here the consumer is fine and the messages are fine, but the service the consumer calls is down. Every message fails, every message gets redelivered, and the consumer hammers the dead service as fast as the broker can hand out work. If there’s a max-attempts setting, it gets used up within minutes and perfectly good messages end up in the dead-letter queue. When the downstream comes back, someone redrives the DLQ and sends the whole backlog at a service that has only just recovered.
During an outage like that, the consumer should notice the dependency is failing and slow down or pause consumption altogether, then pick up again gradually. Retrying each message harder only makes it worse. Treating “this message is broken” and “the dependency is down” as the same kind of failure is exactly how good messages end up in the DLQ.
Redelivery also means a message can be processed more than once, so consumers need their own deduplication. The section on queue consumers in the idempotency article covers the inbox-table approach.
When the retries outlive the outage
There’s one failure pattern that tends to catch teams off guard the first time they see it.
The database has a short problem, say a 20-second failover. Requests slow down, timeouts fire, and clients retry. The extra load makes the database slower, so more requests time out and even more retries go out. Then the failover finishes and the database is healthy again, and the system stays down anyway.
By now incoming traffic is around three times the baseline, because every request is being retried, and the database can’t serve three times its baseline within the clients’ timeouts. So requests keep timing out, which keeps the retry rate high, which keeps the load high. The original trigger is gone, but the loop keeps itself going.
Bronson and colleagues call these metastable failures, in a HotOS paper I’d recommend to anyone who runs a distributed system.[13] What defines them is that the system has a stable bad state: removing the original cause doesn’t bring it back, and something has to break the feedback loop. Retries are one of the sustaining effects the paper discusses, and Google’s SRE chapter on cascading failures describes the same kind of positive feedback, including clients retrying requests that missed their deadlines and making the overload worse.[5]
A few things help break that loop:
- Retry budgets, so the loop can’t grow beyond a small multiple of normal traffic.
- Load shedding on the server. Rejecting early and cheaply, with a clear “overloaded” signal, is much better than accepting work you’re going to time out on anyway. A request that fails in 2ms costs far less than one that fails after holding a connection for 5 seconds.
- A way to turn retries off. In the middle of an incident, being able to drop retries to zero across the fleet with a config change can be what ends it.
- Gradual recovery. When traffic comes back, let it in in stages.
Circuit breakers come up here too, but with a caveat. A client-side breaker that trips on a high error rate can make a partial outage worse, for example when one shard of a dependency is down and the breaker blocks calls to all of them.[2] Brooker’s post also looks at breakers that only stop retries and still let first attempts through, which avoids the worst of that.
The metrics that warn you this is coming are ratios more than counts:
client.retry.ratio retries / first attempts, per dependency
client.retry.budget_exhausted.count
client.deadline_exceeded_before_send.count
server.requests.attempt_number from a header such as X-Retry-Attempt
consumer.redelivery.count
consumer.dlq.count
If clients send their attempt number in a header, the server can see a retry storm building instead of guessing. It’s cheap to add, and it makes the “overloaded, don’t retry” decision much easier to automate.
When I wouldn’t retry at all
Retries aren’t free, and some calls are better off without them:
- Deterministic failures. Validation errors, permission errors, business rule rejections. Retrying them only makes the failure take longer.
- Non-idempotent operations without a key. If you can’t make the retry safe, tell the caller the outcome is unknown. “We couldn’t confirm your payment. Please check your payment history before trying again” is a much better experience than a double charge.
- Calls where the deadline is almost used up. A retry that can’t finish in time adds load and can’t possibly help.
- Interactive requests where the user can retry themselves. A page that fails quickly with a retry button is often better than one that spins for 15 seconds while three attempts happen in the background.
- Anything deep in the call stack. If a layer above already owns retries, adding more here just multiplies them.
Hedging is a related technique worth knowing about. The client sends a second copy of a request if the first hasn’t answered after a short delay, and takes whichever response arrives first. gRPC supports it natively.[14] It can cut tail latency nicely for reads, but it sends duplicates on purpose, so it only belongs on idempotent operations, with a tight cap on how many hedged requests can be in flight.
Failure modes worth testing
Retry behavior is hard to see in unit tests, because a unit test usually covers one client making one call. The tests below need a dependency you control, either a fake or a fault-injecting proxy, and a way to count the requests that reach it.
Downstream returns 503 for everything
Make the dependency return 503 for every request for 60 seconds while normal traffic is running.
Then compare the request rate the dependency sees with the baseline. With a budget in place it should stay fairly close, somewhere around 1.1 to 1.5x. If it’s 3x, you only have per-request limits, and if it’s 9x or 27x, more than one layer is retrying.
Downstream recovers
Keep the same test running and make the dependency healthy again.
Traffic should drop back to baseline quickly, and the success rate should recover with it. If traffic stays high after the dependency is healthy, you’re looking at the beginning of a metastable loop.
Latency just above the timeout
This one is nastier than a plain outage. Have the dependency respond successfully, but a little slower than the client’s timeout, for example a 2-second timeout with responses taking 2.2 seconds.
Every attempt does the full work on the server and counts as a failure on the client. Check how many concurrent
executions of the same operation the server sees, and whether any side effect happens twice. For a POST, this is the
test that tells you whether idempotency and retries were designed together.
Retry-After
Return 429 with Retry-After: 30. The client should wait at least 30 seconds, plus a bit of jitter. Then try the
HTTP-date form, a date in the past, a negative number, an empty value, Retry-After: banana, and
Retry-After: 999999. The client should cope with all of them without retrying immediately or hanging for days.
Non-retryable errors
Return 400, 403, 404 and 422. Each one should result in exactly one request reaching the dependency. A 401
should result in at most two, with a credential refresh in between.
Timeout after the server committed
Let a POST succeed on the server, then drop the connection before the response gets back to the client.
The retry should carry the same idempotency key and the same body, the server should replay the original result, and there should be exactly one resource at the end.
Then kill the client process between attempts and start it again. The retry after the restart should still use the original key.
Deadline nearly spent
Send a request with 50ms of deadline left to a service whose downstream call normally takes 200ms. There should be no downstream call at all, or one at most, and no retries.
Queue consumer during an outage
Take the consumer’s dependency down for five minutes while messages keep flowing.
The consumer should slow down or pause, the messages should stay on the queue, and the dead-letter queue shouldn’t fill up with messages that were never bad in the first place. When the dependency comes back, consumption should pick up again without a burst that knocks it straight back over.
Count the attempts end to end
Make the deepest dependency fail, send one request at the top of the stack, and count how many attempts reach the bottom. Compare that with what you think your configuration allows. If the two numbers don’t match, there’s a retry somewhere that nobody knew about.
Checklist before shipping
- Classify failures before retrying: never processed, processed and rejected, unknown outcome.
- Retry unknown outcomes only for idempotent operations.
- Don’t retry
400,403,404or422by default. - Use exponential backoff with jitter instead of fixed delays.
- Cap retries with a budget as well as a per-request attempt limit.
- Pick one layer to own retries for each hop, and turn retries off in the others.
- Find the retries you didn’t write: SDKs, HTTP clients, drivers, meshes, proxies, brokers.
- Read your retry library’s defaults: how it counts attempts, which exceptions it retries, whether it jitters, and
whether it retries
POST. - In Resilience4j, set the retry predicates explicitly, and ignore
CallNotPermittedExceptionwhen the retry wraps a circuit breaker. - Make inner timeouts shorter than outer ones, and propagate deadlines where you can.
- Skip the retry when the remaining deadline can’t fit another attempt.
- Honor
Retry-After, parse both formats, clamp it, and add jitter. - Generate idempotency keys once per operation, persist them when the operation matters, and reuse them across attempts and restarts.
- Send the same request body on every attempt.
- Cap message redelivery, dead-letter poison messages, and pause consumers when a dependency is down.
- Give operators a way to turn retries down or off during an incident.
- Test sustained failure, recovery, latency just above the timeout, and end-to-end attempt counts.
- Monitor retries as a ratio of first attempts, per dependency.
Every retry runs on the whole fleet
The simple version of retrying asks just one question: did the call fail?
A version that survives production has to ask a few more. What kind of failure was it? Did the server do the work anyway? How much time is left? Is another layer already retrying this? How many retries has this client sent in the last few seconds? Is the dependency asking us to slow down?
Whatever retry loop you write is going to run in every copy of your service, set off by the same outage at the same moment, and that’s the moment worth designing for. A good retry policy hides a flaky network from your users, and when a dependency is really down, it backs off quickly and gives it room to recover.
References
- Marc Brooker, AWS Architecture Blog, “Exponential Backoff And Jitter”.
- Marc Brooker, “Fixing retries with token buckets and circuit breakers”.
- Marc Brooker, “What is Backoff For?”.
- Google SRE Book, Chapter 21: Handling Overload.
- Google SRE Book, Chapter 22: Addressing Cascading Failures.
- gRPC, Retry.
- gRPC, Deadlines.
- AWS SDKs and Tools Reference Guide, Retry behavior.
- RFC 9110, HTTP Semantics, §9.2.2 Idempotent Methods and §10.2.3 Retry-After.
- RFC 9113, HTTP/2, §8.7 Request Reliability.
- RFC 6585, Additional HTTP Status Codes, §4 429 Too Many Requests.
- nginx,
proxy_next_upstream. - Nathan Bronson, Abutalib Aghayev, Aleksey Charapko, Timothy Zhu, “Metastable Failures in Distributed Systems”, HotOS 2021.
- gRPC, Request Hedging.
- Robert M. Metcalfe, David R. Boggs, “Ethernet: Distributed Packet Switching for Local Computer Networks”, Communications of the ACM, 1976.
- Van Jacobson, Michael J. Karels, “Congestion Avoidance and Control”, SIGCOMM 1988.
- Michael T. Nygard, Release It!, Pragmatic Bookshelf (first edition 2007).
- Netflix, Hystrix README: Hystrix Status.
- Finagle, Clients: Retries and RetryBudget.
- Resilience4j, Retry.
- Resilience4j, Spring Boot getting started: Aspect order.
- Spring Framework, Resilience Features:
@Retryable. - Spring Retry, README and
@Retryable. - Polly, Retry resilience strategy.
- Microsoft Learn, Build resilient HTTP apps: standard resilience handler.
- tenacity, Documentation.
- HashiCorp, go-retryablehttp.
- Envoy, Circuit breaker thresholds:
retry_budget. - Resilience4j,
IntervalFunctionsource.