NOTE
1.14 Distributed-System Failures
1. What Is a Distributed-System Failure - A node goes down or the network is unreachable. 2. How to Detect Distributed-System Failures 2.1. Heartbeat Detection 2.2. Gossip Protocol Detection. 3. How to Handle Distributed-System Failures 3.1. Fail-Over - A calls B. If B fails and B has other replicas, A switches to another replica of B.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is a Distributed-System Failure?
A node goes down or the network is unreachable.
2. How to Detect Distributed-System Failures
2.1. Heartbeat Detection
2.2. Gossip Protocol Detection
Distributed Consensus Algorithm: Gossip
3. How to Handle Distributed-System Failures
3.1. Fail-Over
- A calls B. If B fails and B has other replicas, A switches to another replica of B.
3.2. Fail-Fast
- A calls B. If B fails, A immediately throws a failure outward.
3.3. Fail-Safe
- A calls B and C. If A -> B is the primary path while A -> C is a side path, then the failure of A -> C can be ignored.
3.4. Fail-Silent
- A calls B. If B fails, assume B cannot work normally for a period of time. After that, A no longer calls B and directly returns failure.
3.5. Failback
- A calls B. If B fails, restore from B’s backup replica and continue working.
3.6. Forking
- A calls B and calls all replicas of B in parallel. As long as one succeeds, the call succeeds.
3.7. Broadcast
- A calls B and calls all replicas of B in parallel. All of them must succeed for the call to succeed.
4. References
- Distributed High Availability: How to Recover from Failures - Tencent Cloud
- Six Common Fault-Tolerance Mechanisms: Fail-Over, Fail-Fast, Fail-Back, Fail-Safe, Forking and Broadcast - Shoufeng - CNBlogs
- An Article to Thoroughly Clarify failfast, failsafe, failover, failback, and failsilent - whcsrl Technical Network
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub