Skip to content
DevPedia

OperationsGuide 8 of 11

Resilience in distributed systems

How do you keep a partial failure from spreading through the whole system? Resilience mechanisms, when to use them, and where their limits are.

Updated 10 min read

A distributed system changes the failure model completely. Inside a local process, a function either returns a result or throws an exception. In a remote interaction:

the request may never arrive, it may arrive twice, Service B may process the operation and lose the response, Service B may take too long, or only part of the system may be failing.

A partial failure doesn’t stay inside the component that failed. If Payments waits indefinitely on a downed provider, it stops serving checkout; if it retries a charge that already went through, it charges twice. The mechanisms in this article don’t stop things from failing: they limit the damage, make retrying safe, and decide what degrades when not everything can be served.

Timeout

A timeout keeps a dependency from holding resources indefinitely:

But:

A timeout doesn’t mean the operation didn’t happen.

This detail is critical:

The client sees a timeout. The card has already been charged.

Retry

In the face of transient failures, a retry sends the same operation again:

A robust retry usually accounts for exponential backoff, jitter, a maximum number of attempts, and a time budget. Without that ceiling, a slow provider turns into load amplification: every client retries and the provider gets even more traffic.

Retry + non-idempotent operation

This is dangerous:

That’s why timeouts, retries, and idempotency have to be designed together.

Circuit breaker

Lets you stop sending requests to a dependency that is failing:

While it’s OPEN, the system can fail fast instead of piling up timeouts and continuing to hammer the dependency.

Idempotency

Take this:

POST /payments
Idempotency-Key: abc123

On the first request, the provider creates the payment but the response is lost. On the second request, with the same Idempotency-Key, the system recognizes that the operation has already been processed and returns the existing payment instead of creating a new one:

It’s essential for payments, orders, webhooks, and message consumers.

At-least-once delivery

Many messaging systems deliver each message at least once:

That’s why one practical strategy is to combine at-least-once delivery with an idempotent consumer.

Bulkhead

The name comes from the bulkheads of a ship: one flooded compartment shouldn’t sink the rest. In software, you isolate resource pools — connections, threads, semaphores — so a failure in one area doesn’t take another down with it:

If Notifications degrades, it shouldn’t consume every resource available to Payments.

Rate limiting

Controls the incoming flow before the system saturates:

It protects against abusive clients, traffic spikes, and clients that generate extra load without meaning to. Among the best-known implementations are token bucket and leaky bucket. Unlike load shedding, rate limiting acts at the edge: it decides how much to accept, not what to drop once you’re already saturated.

Backpressure

Say a producer runs at 20k msg/s and a consumer at 5k msg/s. The gap builds up:

Backpressure means the consumer or the queue slows the producer down when they can’t keep up, instead of accepting work without limit and letting the queue grow unbounded.

Load shedding

When the system can’t process everything:

It’s better to reject some requests than to let the whole system collapse.

Dead letter queue (DLQ)

A DLQ is the queue messages go to when a consumer couldn’t process them:

The DLQ lets you inspect messages, fix problems, reprocess them later, and keep one bad message from blocking the queue permanently.

Saga

When a business operation spans several services, there’s no single ACID transaction. We can implement a Saga:

If Inventory fails, it triggers a chain of compensating actions:

Graceful degradation

Not every component should carry the same level of criticality:

If Recommendations is down, Checkout stays available and only the Recommendations feature is lost. The architecture keeps what’s critical up.

How to analyze a failure

One practical approach is to walk through a series of questions:

Tools

Depending on the platform: Resilience4j, Envoy, Istio, AWS API Gateway, AWS SQS, Kafka, RabbitMQ, Kubernetes health probes. Tools implement mechanisms. They don’t replace designing for failure modes.

None of these mechanisms is worth anything if you don’t verify that they work. A circuit breaker that has never opened in a controlled environment is a hypothesis, not a protection: it has to be exercised with performance testing under load and with failures injected on purpose.

Share this guide

Search by concept, pattern or practice.