NOTE

6.1 Start with the Hardware

Hardware foundations needed to understand the JMM: von Neumann architecture, caches, coherence, pipelines, reordering, memory barriers, and memory consistency models.

JavaCreated Updated 4 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

To understand the JMM, we first need to understand how the underlying hardware works.

1. Von Neumann Architecture

Von Neumann proposed treating programs as data and storing programs (instructions) and data in the same way. According to this theory, a computer is divided into a controller, arithmetic unit, memory, output devices, and input devices, as shown below.

The arithmetic unit and controller together form the CPU. Whenever the CPU executes instructions or operates on data, it needs to interact with memory. The speed gap between the CPU and memory is enormous, so computer scientists introduced cache between the CPU and main memory to bridge this gap.

2. Cache

After introducing cache, the structure between the CPU and memory is shown below.

2.1. How It Works

With cache, the CPU first attempts to obtain data from registers. If it is not there, it looks in cache. If it is still not there, it looks in main memory. If it is not there either, it looks in secondary storage.

If the data is in secondary storage, it is loaded into main memory, then cache, and finally registers.

Therefore, registers can be considered the cache of cache, cache is the cache of main memory, and main memory is the cache of secondary storage. This is called the memory hierarchy.

2.2. Memory Hierarchy

The closer a storage level is to the top and to the CPU, the faster it is, the smaller its capacity is, and the more expensive it is.

Frequently accessed data can be placed near the top so the CPU can obtain it quickly for computation.

How do we determine which data is frequently accessed? This introduces the principle of locality.

2.3. Principle of Locality

The principle of locality has temporal and spatial aspects.

  • Temporal locality: a memory location that has been referenced may be referenced again.
  • Spatial locality: data near a referenced memory location is likely to be referenced.

According to this principle, the CPU places the data accessed this time and nearby data into the cache layer so it can be accessed quickly next time.

After caches are introduced, CPU access to memory becomes faster, but a cache-coherence problem appears.

3. Cache Coherence / Visibility Problem

Modern CPUs are multi-core processors. If multiple threads concurrently access the same data, each processor’s cache may contain a copy of that data. After processor 1 updates it, when does processor 2 know about the update?

This is the cache-coherence problem: one processor cannot immediately observe data written by another processor.

3.1. How to Solve It

3.1.1. Bus Locking

When a processor reads data from main memory into cache, it locks that data on the bus. Until the processor finishes its operation, other processors cannot read or write that data.

The disadvantage is poor performance. While one processor is reading, other processors cannot perform any operation on that data. Reads could actually proceed concurrently, so cache-coherence protocols such as MESI were introduced.

3.1.2. MESI Cache-Coherence Protocol

Multiple processors can simultaneously read data from main memory into their caches.

When a processor modifies cached data, the coherence mechanism propagates the corresponding state/update.

Other processors observe the state change through cache-coherence communication and invalidate stale cached copies as required.

Next, consider another problem: ordering.

4. CPU Pipeline Technology

CPU instruction execution can roughly be divided into stages such as fetch, decode, computation, memory access, and register write-back.

Like a factory assembly line, when the first instruction reaches the decode stage, the second instruction can simultaneously perform the fetch stage, increasing throughput, as shown below.

Baidu Baike explains CPU pipeline technology as follows:

CPU pipeline technology does not make a single instruction execute faster. Each instruction still needs all of its execution stages. Instead, different stages of multiple instructions execute simultaneously, increasing overall instruction-stream throughput and reducing total program execution time.

4.1. Out-of-Order Execution / Reordering

To better utilize CPU pipelines, the execution order of lines of code in a program may be rearranged by the compiler and CPU according to certain strategies so that instructions can execute in parallel as much as possible. This is called out-of-order execution or reordering.

Reordering can be divided into processor reordering and compiler reordering.

5. Reordering / Ordering Problem

Because modern CPUs are multi-core and use cache layers, data that is logically written later is not necessarily committed to memory last.

In other words, on a multi-core CPU, reordering can cause the final result to differ from what the programmer expected.

5.1. How to Solve It

5.1.1. Use Memory Barriers to Restrict Reordering

Processors of different architectures provide different instructions for issuing memory barriers.

Programming languages expose corresponding mechanisms that map to architecture-specific ordering instructions, such as Java’s volatile.

5.1.1.1. Types of Memory Barriers

Below, Store means writing data and Load means reading data.

Barrier Type Instruction Example Description
LoadLoad Barrier Load1; LoadLoad; Load2 Ensure Load1 is not reordered after Load2
StoreStore Barrier Store1; StoreStore; Store2 Ensure Store1 is not reordered after Store2
LoadStore Barrier Load1; LoadStore; Store2 Ensure Load1 is not reordered after Store2
StoreLoad Barrier Store1; StoreLoad; Load1 Ensure Store1 is not reordered after Load1

The original note groups LoadStore + StoreStore as a release barrier, and LoadLoad + LoadStore as an acquire barrier.

6. Memory Consistency Model

6.1. What Is a Memory Consistency Model?

Baidu Baike explains it as follows:

In essence, it is a contract between software and memory. If software follows the agreed rules, memory can work correctly; otherwise, memory cannot guarantee the correctness of operations.

A memory consistency model is a standard. It specifies which execution orders among all possible execution orders of a program’s memory operations (reads and writes) are correct. Incorrect orders can lead to visibility, ordering, atomicity, and other problems.

Because a memory consistency model is only a standard, different processor architectures can implement it differently. As long as the model is implemented correctly, the required visibility, ordering, and atomicity guarantees can be provided.

6.2. Classification

6.2.1. Sequential Consistency Memory Model

Sequential consistency is an idealized theoretical reference model proposed by computer scientists. It provides programmers with very strong memory-visibility guarantees.

6.2.1.1. Two Main Properties
  1. All operations in one thread must execute according to program order.
  2. Regardless of whether the program is synchronized, all threads observe a single total execution order. In the sequential-consistency model, each operation executes atomically and is immediately visible to all threads.

7. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub