HotShard
Fault Tolerance

Circuit Breaker

A service stops calling a dependency that keeps failing, fails fast instead, and checks now and then whether it has recovered.

When a dependency is slow or down, every call to it waits for a timeout. Each waiting call holds a thread or a connection your service can't use for anything else. Enough of them at once and the caller goes down too. A circuit breaker wraps the call. Once it has seen enough failures, it stops calling and fails immediately. Then it lets a few trial calls through to check whether it is safe to resume.

~6 min read

The problem: one slow dependency takes down the caller#

TL;DRthe 30-second version
  • A dependency that hangs until a timeout holds a thread for the whole wait. Enough of those and the caller has no threads left for healthy work.
  • A circuit breaker watches the outcomes of calls. Once the failure rate is high enough, it stops calling the dependency and fails instantly.
  • Three states: Closed (calls pass), Open (calls rejected fast), Half-Open (a few trial calls test for recovery).
  • It doesn't fix the dependency. It keeps the caller alive and gives the dependency room to recover.

Slowness is worse than plain errors. A dependency that returns an error in 1 ms lets the caller move on. A dependency that hangs until a 10-second timeout holds a thread and a connection for those 10 seconds doing nothing. At even modest request rates, the caller's fixed-size thread pool fills with threads parked on doomed calls. None are left to serve healthy requests.

This is how outages cascadeA slow downstream service makes the caller slow for everyone, including requests that have nothing to do with it. Then the services that depend on the caller slow down too. Retries make it worse: a struggling dependency whose failed calls are each retried 3 times is now handling 3 times the load at its worst moment.

Once you have strong evidence the dependency is unhealthy, the most useful thing you can do is stop calling it. Failing instantly returns the thread, frees the connection, and gives the dependency room to recover. You trade a few requests that might have succeeded for the survival of the whole caller.

The three states#

The breaker sits in front of every call to the dependency. It is always in exactly one of three states.

CLOSEDrequests pass through
failures ≥ threshold →over the window
OPENrequests rejected fast
cooldown elapsed →
HALF-OPENlet 1..k trial requests in
The state machine: a three-state cycle

In Closed, calls pass through. The breaker records each outcome in a sliding window of recent calls. When the window holds enough calls and the failure rate in it crosses a threshold, say 50%, the breaker trips to Open. It needs a rate over a window, not a single failure. A couple of unlucky failures in a mostly healthy window shouldn't trip it. A dependency that is really degrading should, and quickly.

In Open, every call is rejected instantly, with no network call at all. This is the fail-fast part. A cooldown timer starts the moment the breaker opens.

When the cooldown elapses, the breaker doesn't probe on its own. The next call attempt is let through as a trial call. That is Half-Open. If a set number of trial calls succeed in a row, the breaker closes and resets its window. If any trial call fails, it goes back to Open and the cooldown restarts.

PredictThe cooldown just elapsed and 500 requests arrive in the same instant. How many should actually reach the still-unproven dependency?

Hint: What is Half-Open trying to learn, and what happens if all 500 go through?

Only the trial budget, typically 1 to a handful. The other 495 or so are still rejected as if the breaker were Open. Half-Open only needs a tiny sample to decide whether the dependency recovered. If all 500 went through, a still-fragile dependency would be hit with the very surge the breaker exists to prevent, and would likely break again at once. Good implementations cap concurrent Half-Open trials for exactly this reason (resilience4j's permittedNumberOfCallsInHalfOpenState).

How it remembers failures#

  • Consecutive count: one counter of failures in a row. Trip at, say, 5 in a row; any success resets it to 0. Trivial (O(1) state) but blind to mixed traffic. A dependency failing 50% of the time, with successes in between, never trips it.
  • Count-based sliding window: keep the last N outcomes (say a ring buffer of 100 booleans) and trip when the failure rate over them crosses a threshold. O(N) memory, O(1) update. Reacts to the rate, not just streaks.
  • Time-based sliding window: bucket outcomes into the last T seconds (say ten 1-second buckets) and trip on the failure rate inside that window. This bounds how stale the evidence can be, whatever the request rate. Better for bursty or low-traffic services.

Most production breakers also count slow calls as failures. A call that succeeds but takes longer than a slow-call threshold is recorded as a failure. This catches the dependency that isn't erroring, just hanging, before its slowness drains the pool.

Rule of thumbStart with values that need real evidence to trip: a 50% failure rate over a window of at least 20 calls, a 5 to 30 second cooldown, and 1 to 3 Half-Open trials. Then tune against your real traffic and latency. A breaker so twitchy it trips on normal variance turns a resilience tool into an availability bug.

If this comes up in an interview#

The one-linerA circuit breaker watches a dependency's failure rate and, once it trips, rejects calls instantly instead of waiting on timeouts. After a cooldown it lets a few trial calls through, and closes again only when they succeed.

Lead with the failure mode it prevents, which is the caller running out of threads on cascading timeouts. Name the three states and the trigger between each. Interviewers often check whether you know Half-Open is reached lazily, on the next call attempt after the cooldown, and not by a background timer. Then place it in the family: timeouts bound each call, retries with backoff and jitter absorb blips, bulkheads isolate pools, and the breaker decides whether to call at all. Name a real implementation: resilience4j, Polly, or Envoy outlier detection.

What is Half-Open for?

It is the controlled recovery probe. After the cooldown, the breaker can't know whether the dependency recovered without testing, but it must not send full traffic. Half-Open lets a tiny trial budget through. A few successes in a row close the breaker; any failure reopens it. That avoids both staying Open forever and slamming a still-fragile dependency the instant the cooldown ends.

Circuit breaker vs retry. Aren't they the same idea?

Opposite directions. A retry calls more: try again after a failure, for a transient blip. A breaker calls less: stop calling once failures are sustained, for a dependency that is really down. They compose. Retry the occasional blip, but when retries keep failing, the breaker trips and suppresses further attempts so you stop hammering a dead service.

Should a breaker count a 404 or a validation error as a failure?

No. Only count failures that mean the dependency itself is unhealthy: 5xx, timeouts, connection refused. A 404, 400, or 422 is a correct response from a healthy service. Counting them trips the breaker for the wrong reason and breaks healthy traffic.

The trade-offs

A circuit breaker is a bet that recent failures predict the next few. When the bet is right, it saves the system. When it's wrong, it rejects requests that would have succeeded.

You gainYou pay
Fail fast: held threads and connections are released immediately instead of parking on timeouts.False rejections: while Open, you reject requests that might have succeeded if the dependency was only partly degraded.
Protect the dependency: a struggling service gets breathing room instead of a retry pile-on.Tuning burden: threshold, window size, and cooldown must match real traffic and latency, or the breaker misfires.
Contain the blast radius: one unhealthy dependency can't cascade into the whole caller.Reduced visibility: a breaker that trips silently can hide an ongoing outage unless its state changes are monitored and alerted.
Graceful degradation: an Open breaker is a clean hook for a fallback (a cached value, a default, a degraded mode).Complexity: another stateful component with its own failure modes (flapping, herds) to reason about and test.
The core tensionAvailability of the caller versus availability of the feature. An Open breaker keeps the caller healthy by sacrificing the feature that depends on the broken service. That's right when the feature is non-essential or has a fallback. It's wrong when the dependency is critical and a few more attempts might have got through. There is no universal threshold. It's a per-dependency product decision.
The patterns it pairs with, and where it runs
  • Timeout: bounds how long any single call may block. This is the prerequisite. A breaker can't protect you from a slow dependency if individual calls never time out.
  • Retry with exponential backoff and jitter: handles transient blips like one dropped packet or one GC pause. Retries belong inside the breaker, so a tripped breaker suppresses them. Retrying a known-dead dependency is exactly the pile-on you're avoiding.
  • Bulkhead: separate thread pools or concurrency limits per dependency, so a flood toward one can't drain a pool shared with others. The breaker decides whether to call; the bulkhead bounds how many calls can be in flight.
  • Fallback: what you return when the breaker is Open. A cached value, a default, a degraded response, or a fast error.

The other axis is where the breaker lives. In-process libraries (resilience4j on the JVM, Polly on .NET) see real call outcomes, including business versus transport errors and slow calls, and run per dependency. Netflix Hystrix was the implementation that popularized the pattern at scale. Netflix put it in maintenance mode in 2018 and pointed users to resilience4j. A service mesh does it out of process: Envoy's outlier detection ejects upstream hosts that return too many consecutive 5xx responses and re-admits them after a base ejection time, and Istio configures this declaratively. That needs no app code and works across languages, but it only sees what crosses the proxy.

Where it goes wrong
  • Flapping: the breaker trips, recovers on one lucky trial, gets slammed, and trips again, oscillating between Open and Closed instead of settling. Caused by too-short cooldowns or too few Half-Open successes. Fix: require several trial successes in a row and use exponential backoff on the cooldown.
  • Thundering herd on Half-Open: if Half-Open admits all waiting traffic instead of a small trial budget, the first instant after the cooldown re-floods a still-fragile dependency. Fix: cap concurrent Half-Open trials to a handful.
  • No minimum number of calls: on a quiet endpoint the first failure reads as a 100% failure rate and trips the breaker instantly. Require the window to hold a minimum number of calls before a rate can trip it.
  • Wrong failure classification: counting a 404 or a validation 400 as a breaker failure trips the circuit when the dependency is perfectly healthy. Only count 5xx, timeouts, and connection errors.
  • Per-instance state: an in-process breaker is per instance, so 10 caller instances each learn on their own and trip at different times. That's usually fine, but don't assume one global view unless the state is genuinely shared.
References
References

Feedback on this topic →