NOTE
1.1 Message Queue Introduction
What MQ is, queue and publish-subscribe distribution models, push and pull, the benefits and drawbacks of MQ, implementation ideas, and common MQ systems.
This is a historical learning note and may contain outdated or incomplete understanding.
1. What Is MQ?
- A message queue decouples producers and consumers.
- Producers and consumers can actually also be implemented with a database: the producer can write events into the database, and the consumer can periodically poll the database.
- However, this approach puts too much pressure on the database.
2. MQ Message Distribution Models
- There are two message models: the queue model and the publish-subscribe model.
2.1. Queue Model
- One producer -> Queue -> one consumer.
- Messages produced by one producer are independently consumed by one consumer. If multiple consumers consume this Queue at the same time, the data in the Queue will be distributed among these consumers.
- What if I want multiple consumers to each independently consume the complete data in the Queue? I can only create multiple Queues, with each Queue corresponding to one consumer.

2.2. Publish-Subscribe Model
- One producer -> Topic -> multiple consumers.
- Messages produced by one producer can be independently consumed by multiple consumers at the same time.
- If there is only one consumer, then it is the queue model.

3. Message Push Modes
3.1. Push
- The server actively pushes messages to the client.
- Advantages
- Good real-time performance.
- Disadvantages
- The client cannot decide the consumption rate according to its own capability.
- The server needs to maintain a long-lived connection to prevent message loss.
3.2. Pull
- The client pulls messages from the server.
- Advantages
- The client can decide the consumption rate according to its own capability.
- Disadvantages
- Polling puts heavy pressure on the server.
- Real-time performance is a little worse.
4. Benefits of Using MQ
4.1. Decoupling
- Without decoupling, system A needs to directly call the interfaces of other systems to send data. Every time a consuming system is added or removed, system A needs code changes.
- After decoupling, system A only needs to send data to MQ, and other systems can consume data from MQ.

4.2. Asynchronous Processing
- In synchronous-call scenarios, system A needs to call systems B, C, and D. If each takes 200 ms, the total time is 600 ms.
- Using MQ for asynchronous processing can respond to user requests more quickly. System A only needs to write to MQ, taking 100 ms, and the other systems consume and process the data themselves.

4.3. Peak Shaving
- During peak periods, 5K requests hit the database directly and it crashes.
- Use MQ to absorb write requests, and let consumers process them according to their capability.

5. Drawbacks of Using MQ
5.1. Availability Problems
- If MQ goes down, the whole system crashes.
- How to solve it
- Enable cluster mode.
5.2. Message Loss Problems
- There are three reasons why a consumer may not receive a message sent by a producer: the producer failed to send it, MQ lost the message, or the consumer did not consume it successfully but directly reported success to MQ.
- How to solve it
- Producer
- MQ
- Consumer
- Disable automatic commits and use manual commits.
5.3. Idempotency Problems
5.3.1. Producer Sends Repeatedly
Repeated sends by the producer are not a big problem. The main thing is to ensure that consumers do not consume repeatedly.
5.3.2. Consumer Consumes Repeatedly
- The consumer did not commit the offset after consuming (for example, program crash / forced kill / long consumption time / unsubscribe when automatic offset commit is enabled).
- How to solve it
- This is the consumer’s own problem, not an MQ problem, and needs to be handled by the consumer itself.
- Write an idempotent message handler. If writing to a database, first query by the primary key; if it already exists, do not insert it again.
- Track messages and handle duplicates. For example, record the message ID in the database for deduplication, or use a Redis set to store the keys of data that has already been processed.
5.4. Ordering Problems
5.5. Consistency Problems
- A write is considered successful only when A, B, and C all succeed, but if C fails, the data becomes inconsistent.
- How to solve it
- Manual acknowledgement + retry.
6. How to Implement MQ
- A log system can be used to persist messages.
- Producers append messages to the end of the log, while consumers receive messages by reading the log sequentially.
- To ensure high performance, the log can be partitioned, with each partition placed on a different node.
- To ensure high availability, the log can be replicated, with each replica placed on a different node.
- Within each partition, every message has a monotonically increasing sequence number. This ensures that all messages within a partition are fully ordered, while there is no ordering guarantee across different partitions.
7. Common MQ Systems
7.1. RabbitMQ
7.2. Kafka
7.3. Choosing an MQ
| ActiveMQ | RabbitMQ | RocketMQ | Kafka | |
|---|---|---|---|---|
| Features | Incomplete, messages may be lost | Complete (messages can be routed in various ways) | Complete | Incomplete (ordinary message sending and receiving), messages may be lost |
| QPS | Tens of thousands | Tens of thousands | Hundreds of thousands | Hundreds of thousands |
| Community activity | Inactive | Active | Active | Active |
| Language | Java | Erlang | Java | Java |
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub