NOTE

1.14 Distributed-System Failures

1. What Is a Distributed-System Failure - A node goes down or the network is unreachable. 2. How to Detect Distributed-System Failures 2.1. Heartbeat Detection 2.2. Gossip Protocol Detection. 3. How to Handle Distributed-System Failures 3.1. Fail-Over - A calls B. If B fails and B has other replicas, A switches to another replica of B.

Distributed SystemsCreated Updated 1 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is a Distributed-System Failure?

A node goes down or the network is unreachable.

2. How to Detect Distributed-System Failures

2.1. Heartbeat Detection

2.2. Gossip Protocol Detection

Distributed Consensus Algorithm: Gossip

3. How to Handle Distributed-System Failures

3.1. Fail-Over

  • A calls B. If B fails and B has other replicas, A switches to another replica of B.

3.2. Fail-Fast

  • A calls B. If B fails, A immediately throws a failure outward.

3.3. Fail-Safe

  • A calls B and C. If A -> B is the primary path while A -> C is a side path, then the failure of A -> C can be ignored.

3.4. Fail-Silent

  • A calls B. If B fails, assume B cannot work normally for a period of time. After that, A no longer calls B and directly returns failure.

3.5. Failback

  • A calls B. If B fails, restore from B’s backup replica and continue working.

3.6. Forking

  • A calls B and calls all replicas of B in parallel. As long as one succeeds, the call succeeds.

3.7. Broadcast

  • A calls B and calls all replicas of B in parallel. All of them must succeed for the call to succeed.

4. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub