NOTE
How to Design a Fault-Tolerance Component
What fault tolerance is, service avalanche, and common techniques including timeouts, retries, rate limiting, circuit breaking, isolation, and degradation.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is Fault Tolerance
Under a microservice architecture, if one service fails, it can affect the entire call chain and make the overall service unavailable.
2. Why Fault Tolerance Is Needed
Ensure service availability and prevent the avalanche effect.
2.1. Service Avalanche
When multiple microservices call one another, suppose Service A calls Services B and C, Service B calls Services D and E, and Service C calls Services F and G… This is called fan-out.
The phenomenon in which “service provider unavailability” (cause) leads to “service caller unavailability” (result), and the unavailability gradually amplifies. In other words, one service on a fan-out chain fails and causes all services on the entire chain to fail.
As shown in the figure, Service C calls Service D. If Service D has a problem and never returns (response time is too long or the service is unavailable), Service C continues to call Service D, and Service D still does not return… Retries continue until Service C’s thread resources are exhausted and it can no longer provide other services. Services A and B that call other interfaces of Service C are therefore also affected, eventually affecting the entire microservice system.

3. How to Design a Fault-Tolerance Component
3.1. Timeout

- To prevent Service C from requesting Service D and never receiving a response, causing resources such as Service C’s threads to remain held and unable to be released, a timeout needs to be configured.
- HTTP and RPC frameworks generally have timeout settings, such as context.md.
3.2. Retry

- If Service C encounters a network error when requesting Service D, a timeout is triggered. But the network problem may be temporary, so retrying a few times may solve it.
- For example, retry.md.
3.3. Rate Limiting

- If Service C accesses a database that supports at most 2,000 concurrent requests per second, and Services A and B each send 2,000 concurrent requests to it, Service C will fail. Therefore, C’s concurrency needs to be limited and excess requests rejected.
- How to Design a Rate-Limiting System.md
3.4. Circuit Breaking

- When the error count reaches a threshold, stop calling the target module. Resume calls when the situation improves.
- When a downstream service suddenly becomes unavailable or responds too slowly for some reason, the upstream service stops calling the target service and returns directly in order to preserve its own overall availability and release resources quickly.
- If the target service improves, calls resume.
- This is the circuit-breaker pattern.
3.5. Isolation

- Service C requests Services D and E. If Service D fails, Service C keeps timing out and retrying, so resources such as Service C’s threads are all held by Service D and cannot be used to request Service E. Therefore, the resources for C -> D and C -> E need to be isolated.
- This is the bulkhead-isolation pattern.
3.6. Degradation

- When an upstream service calls a downstream service and the downstream service is unavailable or responds too slowly, execute the upstream service’s local fallback plan.
- Custom handling.
- fail-fast.
- fail-silent.
- Service C calls Service D. If Service D fails, Service C’s local cache can be used.

Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub