NOTE

4.3 Distributed-System Replication Architecture: Leader-Follower Replication

1. What Is Leader-Follower - There is exactly one Leader among the replicas, and all others are Followers. 2. Leader-Follower Use Cases - A single data center. 3. Leader Election - Select one replica as the Leader and use the other replicas as Followers.

Distributed SystemsCreated Updated 2 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is Leader-Follower?

  • There is exactly one Leader among the replicas, and all others are Followers.

2. Leader-Follower Use Cases

  • A single data center.

3. Leader Election

Select one replica as the Leader and use the other replicas as Followers.

3.1. Election Methods

  1. Manual
    • Manually designate one node as Leader.
  2. Depend on a centralized component
    • For example, ZooKeeper.
  3. Election algorithms

3.2. Split-Brain Problem

If a network partition occurs, each partition has one Leader. This is the split-brain problem.

Solution: introduce a majority mechanism.

4. How Leader-Follower Works

4.1. Leader Election at Startup

  1. Manual.
  2. Centralized.
  3. Election algorithm.
    • A newly started cluster has no data, so the Leader is determined according to ID size.
    • A starts. Through ping, it obtains the node list (A), elects the node with the smallest ID (only itself) as master, but the minimum of two nodes is not satisfied, so it loops and waits.
    • B starts. Through ping, it obtains the node list (A, B), elects node A with the smallest ID as master, satisfies the minimum of two nodes, and forms a cluster.
    • C starts. Through ping, it obtains the node list (A, B, C). Because there is already a master, it directly joins this cluster.

4.2. Data Synchronization

4.2.1. Synchronization Process

  • When a Follower connects to the Leader for the first time, it needs to synchronize all of the Leader’s data. This process is called full synchronization. The process is:
    1. The Leader creates a snapshot of the data at the current moment.
    2. The Leader sends the snapshot to the new Follower.
    3. The Leader continues serving client writes.
    4. The Follower replays the snapshot.
    5. The Follower pulls all data changes after the Leader’s snapshot.

4.2.2. Synchronization Method

4.2.3. Synchronization Log

4.3. Request Processing

Who handles it and how it is routed.

4.3.1. Read Requests

Can be handled by the Leader or a Follower.

4.3.2. Write Requests

Must be handled by the Leader.

If the request is routed to the Leader, the Leader processes it and then synchronizes it to Followers.

If the request is routed to a Follower, the Follower needs to forward it to the Leader, or the Follower tells the client to redirect to the Leader. After the Leader processes it, it synchronizes it to Followers.

4.4. Failure Handling

4.4.1. Failure Detection

Distributed-System Failures

4.4.2. Failure Recovery

4.4.2.1. Follower Crash
  • After a Follower crashes and restarts, it can know from its local log where replication had reached. After reconnecting to the Leader, it only needs to replicate from that position onward. This is called incremental synchronization. The process is:
    1. The Follower restarts and connects to the Leader.
    2. The Follower reads the local log position and pulls data changes after that position from the Leader.
    3. The Follower replays these data changes.
4.4.2.2. Leader Crash
  • After the Leader crashes:
    • Select one Follower and promote it to Leader.
    • Inform the client and the other Followers that the Leader has changed.

5. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub