NOTE

1.1 Message Queue Introduction

What MQ is, queue and publish-subscribe distribution models, push and pull, the benefits and drawbacks of MQ, implementation ideas, and common MQ systems.

Message QueuesCreated Updated 3 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is MQ?

  • A message queue decouples producers and consumers.
  • Producers and consumers can actually also be implemented with a database: the producer can write events into the database, and the consumer can periodically poll the database.
  • However, this approach puts too much pressure on the database.

2. MQ Message Distribution Models

  • There are two message models: the queue model and the publish-subscribe model.

2.1. Queue Model

  • One producer -> Queue -> one consumer.
    • Messages produced by one producer are independently consumed by one consumer. If multiple consumers consume this Queue at the same time, the data in the Queue will be distributed among these consumers.
    • What if I want multiple consumers to each independently consume the complete data in the Queue? I can only create multiple Queues, with each Queue corresponding to one consumer.
  • Message Distribution Model - Queue Model

2.2. Publish-Subscribe Model

  • One producer -> Topic -> multiple consumers.
    • Messages produced by one producer can be independently consumed by multiple consumers at the same time.
    • If there is only one consumer, then it is the queue model.
  • Message Distribution Model - Topic Model

3. Message Push Modes

3.1. Push

  • The server actively pushes messages to the client.
  • Advantages
    • Good real-time performance.
  • Disadvantages
    • The client cannot decide the consumption rate according to its own capability.
    • The server needs to maintain a long-lived connection to prevent message loss.

3.2. Pull

  • The client pulls messages from the server.
  • Advantages
    • The client can decide the consumption rate according to its own capability.
  • Disadvantages
    • Polling puts heavy pressure on the server.
    • Real-time performance is a little worse.

4. Benefits of Using MQ

4.1. Decoupling

  • Without decoupling, system A needs to directly call the interfaces of other systems to send data. Every time a consuming system is added or removed, system A needs code changes.
  • After decoupling, system A only needs to send data to MQ, and other systems can consume data from MQ. Benefits of Message Queues - Decoupling

4.2. Asynchronous Processing

  • In synchronous-call scenarios, system A needs to call systems B, C, and D. If each takes 200 ms, the total time is 600 ms.
  • Using MQ for asynchronous processing can respond to user requests more quickly. System A only needs to write to MQ, taking 100 ms, and the other systems consume and process the data themselves.

Benefits of Message Queues - Asynchronous Processing

4.3. Peak Shaving

  • During peak periods, 5K requests hit the database directly and it crashes.
  • Use MQ to absorb write requests, and let consumers process them according to their capability.

Unnamed Drawing - Peak Shaving

5. Drawbacks of Using MQ

5.1. Availability Problems

5.2. Message Loss Problems

5.3. Idempotency Problems

5.3.1. Producer Sends Repeatedly

Repeated sends by the producer are not a big problem. The main thing is to ensure that consumers do not consume repeatedly.

5.3.2. Consumer Consumes Repeatedly

  • The consumer did not commit the offset after consuming (for example, program crash / forced kill / long consumption time / unsubscribe when automatic offset commit is enabled).
  • How to solve it
    • This is the consumer’s own problem, not an MQ problem, and needs to be handled by the consumer itself.
    • Write an idempotent message handler. If writing to a database, first query by the primary key; if it already exists, do not insert it again.
    • Track messages and handle duplicates. For example, record the message ID in the database for deduplication, or use a Redis set to store the keys of data that has already been processed.

5.4. Ordering Problems

5.5. Consistency Problems

  • A write is considered successful only when A, B, and C all succeed, but if C fails, the data becomes inconsistent.
  • How to solve it
    • Manual acknowledgement + retry.

6. How to Implement MQ

  • A log system can be used to persist messages.
  • Producers append messages to the end of the log, while consumers receive messages by reading the log sequentially.
  • To ensure high performance, the log can be partitioned, with each partition placed on a different node.
  • To ensure high availability, the log can be replicated, with each replica placed on a different node.
  • Within each partition, every message has a monotonically increasing sequence number. This ensures that all messages within a partition are fully ordered, while there is no ordering guarantee across different partitions.

7. Common MQ Systems

7.1. RabbitMQ

7.2. Kafka

7.3. Choosing an MQ

ActiveMQ RabbitMQ RocketMQ Kafka
Features Incomplete, messages may be lost Complete (messages can be routed in various ways) Complete Incomplete (ordinary message sending and receiving), messages may be lost
QPS Tens of thousands Tens of thousands Hundreds of thousands Hundreds of thousands
Community activity Inactive Active Active Active
Language Java Erlang Java Java

8. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub