NOTE
6.1 Start with the Hardware
Hardware foundations needed to understand the JMM: von Neumann architecture, caches, coherence, pipelines, reordering, memory barriers, and memory consistency models.
This is a historical learning note and may contain outdated or incomplete understanding.
To understand the JMM, we first need to understand how the underlying hardware works.
1. Von Neumann Architecture
Von Neumann proposed treating programs as data and storing programs (instructions) and data in the same way. According to this theory, a computer is divided into a controller, arithmetic unit, memory, output devices, and input devices, as shown below.

The arithmetic unit and controller together form the CPU. Whenever the CPU executes instructions or operates on data, it needs to interact with memory. The speed gap between the CPU and memory is enormous, so computer scientists introduced cache between the CPU and main memory to bridge this gap.
2. Cache
After introducing cache, the structure between the CPU and memory is shown below.

2.1. How It Works
With cache, the CPU first attempts to obtain data from registers. If it is not there, it looks in cache. If it is still not there, it looks in main memory. If it is not there either, it looks in secondary storage.
If the data is in secondary storage, it is loaded into main memory, then cache, and finally registers.
Therefore, registers can be considered the cache of cache, cache is the cache of main memory, and main memory is the cache of secondary storage. This is called the memory hierarchy.
2.2. Memory Hierarchy
The closer a storage level is to the top and to the CPU, the faster it is, the smaller its capacity is, and the more expensive it is.
Frequently accessed data can be placed near the top so the CPU can obtain it quickly for computation.
How do we determine which data is frequently accessed? This introduces the principle of locality.
2.3. Principle of Locality
The principle of locality has temporal and spatial aspects.
- Temporal locality: a memory location that has been referenced may be referenced again.
- Spatial locality: data near a referenced memory location is likely to be referenced.
According to this principle, the CPU places the data accessed this time and nearby data into the cache layer so it can be accessed quickly next time.
After caches are introduced, CPU access to memory becomes faster, but a cache-coherence problem appears.
3. Cache Coherence / Visibility Problem
Modern CPUs are multi-core processors. If multiple threads concurrently access the same data, each processor’s cache may contain a copy of that data. After processor 1 updates it, when does processor 2 know about the update?
This is the cache-coherence problem: one processor cannot immediately observe data written by another processor.

3.1. How to Solve It
3.1.1. Bus Locking
When a processor reads data from main memory into cache, it locks that data on the bus. Until the processor finishes its operation, other processors cannot read or write that data.
The disadvantage is poor performance. While one processor is reading, other processors cannot perform any operation on that data. Reads could actually proceed concurrently, so cache-coherence protocols such as MESI were introduced.
3.1.2. MESI Cache-Coherence Protocol
Multiple processors can simultaneously read data from main memory into their caches.
When a processor modifies cached data, the coherence mechanism propagates the corresponding state/update.
Other processors observe the state change through cache-coherence communication and invalidate stale cached copies as required.
Next, consider another problem: ordering.
4. CPU Pipeline Technology
CPU instruction execution can roughly be divided into stages such as fetch, decode, computation, memory access, and register write-back.
Like a factory assembly line, when the first instruction reaches the decode stage, the second instruction can simultaneously perform the fetch stage, increasing throughput, as shown below.

Baidu Baike explains CPU pipeline technology as follows:
CPU pipeline technology does not make a single instruction execute faster. Each instruction still needs all of its execution stages. Instead, different stages of multiple instructions execute simultaneously, increasing overall instruction-stream throughput and reducing total program execution time.
4.1. Out-of-Order Execution / Reordering
To better utilize CPU pipelines, the execution order of lines of code in a program may be rearranged by the compiler and CPU according to certain strategies so that instructions can execute in parallel as much as possible. This is called out-of-order execution or reordering.
Reordering can be divided into processor reordering and compiler reordering.
5. Reordering / Ordering Problem
Because modern CPUs are multi-core and use cache layers, data that is logically written later is not necessarily committed to memory last.
In other words, on a multi-core CPU, reordering can cause the final result to differ from what the programmer expected.
5.1. How to Solve It
5.1.1. Use Memory Barriers to Restrict Reordering
Processors of different architectures provide different instructions for issuing memory barriers.
Programming languages expose corresponding mechanisms that map to architecture-specific ordering instructions, such as Java’s volatile.
5.1.1.1. Types of Memory Barriers
Below, Store means writing data and Load means reading data.
| Barrier Type | Instruction Example | Description |
|---|---|---|
| LoadLoad Barrier | Load1; LoadLoad; Load2 | Ensure Load1 is not reordered after Load2 |
| StoreStore Barrier | Store1; StoreStore; Store2 | Ensure Store1 is not reordered after Store2 |
| LoadStore Barrier | Load1; LoadStore; Store2 | Ensure Load1 is not reordered after Store2 |
| StoreLoad Barrier | Store1; StoreLoad; Load1 | Ensure Store1 is not reordered after Load1 |
The original note groups LoadStore + StoreStore as a release barrier, and LoadLoad + LoadStore as an acquire barrier.
6. Memory Consistency Model
6.1. What Is a Memory Consistency Model?
Baidu Baike explains it as follows:
In essence, it is a contract between software and memory. If software follows the agreed rules, memory can work correctly; otherwise, memory cannot guarantee the correctness of operations.
A memory consistency model is a standard. It specifies which execution orders among all possible execution orders of a program’s memory operations (reads and writes) are correct. Incorrect orders can lead to visibility, ordering, atomicity, and other problems.
Because a memory consistency model is only a standard, different processor architectures can implement it differently. As long as the model is implemented correctly, the required visibility, ordering, and atomicity guarantees can be provided.
6.2. Classification
6.2.1. Sequential Consistency Memory Model
Sequential consistency is an idealized theoretical reference model proposed by computer scientists. It provides programmers with very strong memory-visibility guarantees.
6.2.1.1. Two Main Properties
- All operations in one thread must execute according to program order.
- Regardless of whether the program is synchronized, all threads observe a single total execution order. In the sequential-consistency model, each operation executes atomically and is immediately visible to all threads.
7. References
- Discussion of Out-of-Order Execution and Memory Barriers - CSDN
- CPU Pipeline - Baidu Baike
- How Should the Principle of Locality in Computer Systems Be Understood? - Zhihu
- Why Does a Longer CPU Pipeline Make a Processor Faster? - Zhihu
- Memory Consistency Model - Baidu Baike
- Memory Consistency Model - Wikipedia
- Memory model (programming) - Wikipedia
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub