Retries with exponential backoff and jitter

ยท 2 min read

Retrying a failed call is easy. Retrying it without making the outage worse takes backoff, jitter and some care about which requests are safe to repeat.

Any integration that talks to another service over a network will see transient failures: a timeout, a dropped connection, a 503 during a deployment. Retrying is the right response. Retrying badly can turn a two-second blip into a ten-minute outage.

Three ways to get retries wrong

Retrying immediately. If the other service is overloaded, hitting it again straight away adds load at the worst possible moment.

Retrying on a fixed interval. Better, but when hundreds of clients failed at the same instant, they all retry at the same instant too. The dependency recovers, gets hit by a synchronized wave, and falls over again. This is the "thundering herd".

Retrying everything. A 400 Bad Request will fail the same way every time. Retrying it just wastes time and quota.

Backoff and jitter

Exponential backoff doubles the wait after each failure: 1 second, 2, 4, 8. The dependency gets more breathing room the longer it struggles.

Jitter adds randomness to each delay, so clients that failed together spread their retries out instead of moving in lockstep. Backoff without jitter still produces waves; jitter breaks them up.

In .NET: let the resilience pipeline do it

You don't need to hand-write retry loops. The Microsoft.Extensions.Http.Resilience package (built on Polly) plugs straight into IHttpClientFactory. The quickest start is the standard handler, which already includes retries with exponential backoff and jitter, plus timeouts and a circuit breaker:

builder.Services.AddHttpClient<InventoryClient>(c =>
        c.BaseAddress = new Uri("https://inventory.example.com/"))
    .AddStandardResilienceHandler();

If you want to tune the retry yourself:

builder.Services.AddHttpClient<InventoryClient>(c =>
        c.BaseAddress = new Uri("https://inventory.example.com/"))
    .AddResilienceHandler("inventory", pipeline =>
    {
        pipeline.AddRetry(new HttpRetryStrategyOptions
        {
            MaxRetryAttempts = 4,
            BackoffType = DelayBackoffType.Exponential,
            UseJitter = true,
            Delay = TimeSpan.FromSeconds(1)
        });
    });

The HTTP retry options already treat timeouts, connection failures, 408, 429 and 5xx responses as transient, and leave other 4xx errors alone.

Only retry what's safe to repeat

A GET can be retried freely. A POST that creates an order might have succeeded on the server even though your client saw a timeout. Retrying it creates a second order.

You have two options:

  • Turn off retries for unsafe methods. With the standard handler, options.Retry.DisableForUnsafeHttpMethods() does exactly that.
  • Make the endpoint idempotent, usually with an idempotency key the server uses to recognize a repeated request. Many payment and messaging APIs support this.

Put a ceiling on it

Retries multiply latency. Four attempts with backoff can easily add 15 seconds before the caller sees an error. Pair retries with an overall timeout so the whole operation gives up in a predictable time, and make sure that timeout is shorter than whatever is waiting on you.

Takeaway

Retry transient failures, not every failure. Back off exponentially, add jitter so clients don't retry in sync, only repeat requests that are safe to repeat, and cap the total time. The resilience handler for HttpClient gives you all of this with a few lines of configuration.