Circuit breakers stop you hammering a failing service
Retries help with brief failures. When a dependency is down for minutes, retries just pile on. A circuit breaker notices and fails fast until it recovers.
Retries are great for a dropped connection or a brief 503. They're the wrong tool when a dependency is properly down. Every request still waits for timeouts and burns through its retries, your threads and connections pile up, your own response times collapse, and the struggling service gets even more traffic while it tries to recover.
A circuit breaker handles that case. It watches the results of calls to a dependency, and when failures pass a threshold it stops calling it for a while.
The three states
- Closed. Normal operation. Calls go through and the breaker tracks how many fail.
- Open. Too many recent failures. Calls fail immediately without touching the network, for a set break duration.
- Half-open. After the break, a trial call is allowed through. If it succeeds, the breaker closes. If it fails, it opens again.
Failing fast sounds bad, but it's far better than every request waiting 30 seconds to fail anyway. Your service stays responsive, and the dependency gets room to recover.
In .NET
The standard resilience handler for HttpClient already includes a circuit breaker. To set your own thresholds:
builder.Services.AddHttpClient<PaymentsClient>()
.AddResilienceHandler("payments", pipeline =>
{
pipeline.AddRetry(new HttpRetryStrategyOptions
{
MaxRetryAttempts = 2,
BackoffType = DelayBackoffType.Exponential,
UseJitter = true
});
pipeline.AddCircuitBreaker(new HttpCircuitBreakerStrategyOptions
{
FailureRatio = 0.5, // open when half the calls fail...
MinimumThroughput = 20, // ...out of at least 20 calls...
SamplingDuration = TimeSpan.FromSeconds(30), // ...within 30 seconds
BreakDuration = TimeSpan.FromSeconds(15) // then fail fast for 15 seconds
});
pipeline.AddTimeout(TimeSpan.FromSeconds(5));
});
Strategies run in the order you add them, from the outside in. Here, each attempt has a 5-second timeout, the circuit breaker sees every attempt, and the retry wraps both. When the circuit is open, calls throw BrokenCircuitException straight away.
Handle the open circuit
Decide what your code does when the breaker is open, rather than letting the exception bubble up as a generic 500:
try
{
return await payments.GetStatusAsync(paymentId, ct);
}
catch (BrokenCircuitException)
{
logger.LogWarning("Payments service unavailable, returning cached status for {PaymentId}", paymentId);
return await cache.GetLastKnownStatusAsync(paymentId, ct);
}
Depending on the case, fall back to cached data, queue the work for later, or return a clear 503 to your own caller.
Choosing thresholds
- Set a minimum throughput, or a couple of failures at 3 a.m. with almost no traffic will trip the breaker.
- Keep the break short, seconds rather than minutes. The half-open state reopens it quickly if the dependency is still down.
- One breaker per dependency. If a client calls several hosts,
SelectPipelineByAuthority()gives each host its own breaker, so one failing host doesn't block the others.
Takeaway
Retries handle blips; circuit breakers handle outages. Put a breaker in front of each external dependency, make sure it needs real traffic before it trips, and decide what your code does when the circuit is open.