NOTE

2.4 Kafka Architecture

Kafka topology, producers, brokers, consumers, replication, failure handling, and partitioning.

Message QueuesCreated Updated 2 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. Kafka Topology

Kafka - Architecture A Producer sends messages to a specific Topic. Messages in the Topic are stored on Brokers, and a Consumer consumes messages by subscribing to a specific Topic.

2. Three Major Components

2.1. Producer

Kafka Producer

2.2. Broker

Kafka Broker Kafka Topic

2.3. Consumer

Kafka Consumer

3. Distributed Principles

3.1. Primary-Replica Replication

3.1.1. Leader Election

One replica is selected as the Leader, and the other replicas act as Followers.

3.1.1.1. Election Method

The Controller node among Brokers is elected by ZooKeeper. The Leader in a Partition is elected by the Controller.

3.1.2. Data Synchronization

3.1.2.1. Synchronization Process

  • When a follower connects to a leader for the first time, it needs to synchronize all of the leader’s data. This process is called full synchronization. The process is as follows:
    1. The leader creates a snapshot of the data at the current moment.
    2. The leader sends the snapshot to the new follower.
    3. The leader continues serving client writes.
    4. The follower replays the snapshot.
    5. The follower pulls all data changes after the leader’s snapshot.

3.1.2.2. Synchronization Method

Kafka Broker has an ack parameter, which can be regarded as supporting synchronous and asynchronous modes.

3.1.2.3. Synchronized Log

3.1.3. Request Processing

3.1.3.1. Read Requests

They must be handled by the leader.

3.1.3.2. Write Requests

They must be handled by the leader. If a request is routed to the leader, the leader processes it and synchronizes it to the followers.

3.1.4. Failure Handling

3.1.4.1. Failure Detection

The Controller is highly available based on ZooKeeper. Leader

3.1.4.2. Failure Recovery

3.1.4.2.1. Follower Goes Down
  • After a Follower goes down and restarts, it can determine from its local log where replication had reached. After reconnecting to the Leader, it can continue replicating from that position. This is called incremental synchronization. The process is as follows:
    1. The follower restarts and connects to the leader.
    2. The follower reads the position in the local log and pulls data changes after that position from the leader.
    3. The follower replays these data changes.
3.1.4.2.2. Leader Goes Down
  • After the Leader goes down, it is necessary to:
    • select a follower and promote it to leader;
    • inform the clients and other followers that the leader has changed.

3.2. Partitioning

3.2.1. Split Data

3.2.1.1. Choosing the Split Key

Use a key.

3.2.1.2. Data Splitting Strategy

3.2.2. Partition Assignment

Dynamic assignment. Automatic partition rebalancing.

3.2.3. Request Processing

3.2.3.1. Routing Component

3.2.3.2. Routing Process

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub