NOTE
2.4 Kafka Architecture
Kafka topology, producers, brokers, consumers, replication, failure handling, and partitioning.
This is a historical learning note and may contain outdated or incomplete understanding.
1. Kafka Topology
A Producer sends messages to a specific Topic. Messages in the Topic are stored on Brokers, and a Consumer consumes messages by subscribing to a specific Topic.
2. Three Major Components
2.1. Producer
2.2. Broker
2.3. Consumer
3. Distributed Principles
3.1. Primary-Replica Replication
3.1.1. Leader Election
One replica is selected as the Leader, and the other replicas act as Followers.
3.1.1.1. Election Method
The Controller node among Brokers is elected by ZooKeeper. The Leader in a Partition is elected by the Controller.
3.1.2. Data Synchronization
3.1.2.1. Synchronization Process
- When a follower connects to a leader for the first time, it needs to synchronize all of the leader’s data. This process is called full synchronization. The process is as follows:
- The leader creates a snapshot of the data at the current moment.
- The leader sends the snapshot to the new follower.
- The leader continues serving client writes.
- The follower replays the snapshot.
- The follower pulls all data changes after the leader’s snapshot.
3.1.2.2. Synchronization Method
Kafka Broker has an ack parameter, which can be regarded as supporting synchronous and asynchronous modes.
3.1.2.3. Synchronized Log
3.1.3. Request Processing
3.1.3.1. Read Requests
They must be handled by the leader.
3.1.3.2. Write Requests
They must be handled by the leader. If a request is routed to the leader, the leader processes it and synchronizes it to the followers.
3.1.4. Failure Handling
3.1.4.1. Failure Detection
The Controller is highly available based on ZooKeeper. Leader
3.1.4.2. Failure Recovery
3.1.4.2.1. Follower Goes Down
- After a Follower goes down and restarts, it can determine from its local log where replication had reached. After reconnecting to the Leader, it can continue replicating from that position. This is called incremental synchronization. The process is as follows:
- The follower restarts and connects to the leader.
- The follower reads the position in the local log and pulls data changes after that position from the leader.
- The follower replays these data changes.
3.1.4.2.2. Leader Goes Down
- After the Leader goes down, it is necessary to:
- select a follower and promote it to leader;
- inform the clients and other followers that the leader has changed.
3.2. Partitioning
3.2.1. Split Data
3.2.1.1. Choosing the Split Key
Use a key.
3.2.1.2. Data Splitting Strategy
3.2.2. Partition Assignment
Dynamic assignment. Automatic partition rebalancing.
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub