Concurrency Programming (8): Read-Write Locks — From Language Rules to the CPU

Uses the same counter and counter + ready examples to trace read-write locks from Java language-level rules through the JDK, HotSpot, and x86-64, explaining Atomicity, Visibility, Ordering, and the boundary with Mutex.

EnglishPublished Updated 10/04/20268 min read
Concurrency Programming (8): Read-Write Locks — From Language Rules to the CPU

Table of Contents


0. What Does This Article Answer Next?

The previous article ended by returning to Mutex: it protects an entire critical section, but it allows only one execution unit to enter at a time, regardless of whether that unit reads or writes.

If two threads only read shared state, they do not need to block each other:

Thread A                         Thread B

read counter                    read counter

A read-write lock refines Mutex by distinguishing Readers from Writers:

Read  + Read   → concurrent
Read  + Write  → exclusive
Write + Write  → exclusive

This article continues with the same two examples:

  • counter++ to see how the Write Lock provides Atomicity;
  • counter + ready to see how Write Lock → Read Lock establishes Visibility and Ordering.

As before, we start with the language-level rules, then trace how the JDK, HotSpot, and the CPU implement them.


1. What Does Java ReadWriteLock Guarantee at the Language Level?

1.1 Atomicity

Java’s ReadWriteLock defines the Reader / Writer boundary directly:

“The read lock may be held simultaneously by multiple reader threads, so long as there are no writers.”

It also states:

“The write lock is exclusive.”

Reads may therefore proceed concurrently, while writes remain exclusive.

Use the same counter++ example:

ReentrantReadWriteLock rw = new ReentrantReadWriteLock();
Lock writeLock = rw.writeLock();

writeLock.lock();
try {
    counter++;
} finally {
    writeLock.unlock();
}

With two threads updating concurrently:

Initial: counter = 0

Thread A                         Thread B

writeLock.lock()
                                 writeLock.lock()
                                 wait
read counter -> 0
counter + 1 -> 1
write counter -> 1
writeLock.unlock()
                                 writeLock.lock() returns
                                 read counter -> 1
                                 counter + 1 -> 2
                                 write counter -> 2
                                 writeLock.unlock()

The final value is:

counter = 2

As with Mutex, Atomicity comes from serializing conflicting critical sections under the Write Lock, not from turning counter++ itself into an atomic instruction.

A Read Lock should protect reads only. Multiple Readers may enter together, so it cannot be used to protect shared writes.

1.2 Visibility and Ordering

ReadWriteLock also defines memory synchronization between the Write Lock and the associated Read Lock:

“A thread successfully acquiring the read lock will see all updates made upon previous release of the write lock.”

Use the same counter / ready example:

int counter = 0;
boolean ready = false;

ReentrantReadWriteLock rw = new ReentrantReadWriteLock();
Lock readLock = rw.readLock();
Lock writeLock = rw.writeLock();

// Thread A
writeLock.lock();
try {
    counter = 1;
    ready = true;
} finally {
    writeLock.unlock();
}

// Thread B
readLock.lock();
try {
    if (ready) {
        System.out.println(counter);
    }
} finally {
    readLock.unlock();
}

If B successfully acquires the Read Lock after A releases the Write Lock:

Thread A                              Thread B

counter = 1
    │
    │ program order
    ▼
ready = true
    │
    ▼
writeLock.unlock()
    │
    │ memory synchronization
    ▼
                                    readLock.lock() returns
                                        │
                                        │ program order
                                        ▼
                                    read ready == true
                                    read counter

By happens-before:

write counter
    ↓ happens-before
write ready
    ↓
writeLock.unlock
    ↓ memory synchronization
readLock.lock returns
    ↓ happens-before
read counter

So once B has acquired the subsequent Read Lock and observes ready == true, the earlier counter = 1 must be visible. This outcome is not allowed:

ready == true
counter == 0

That gives us two guarantees:

  • Visibility: writes completed by the Writer before releasing the Write Lock are visible to a later Reader that acquires the Read Lock;
  • Ordering: the Writer’s writes, Write Lock release, Read Lock acquisition, and the Reader’s reads form a happens-before order.

2. How Does the JDK Implement ReentrantReadWriteLock?

2.1 Layers

Start with the implementation stack:

Java Source Code
Example: readLock.lock() / writeLock.lock()
Role: declare read and write critical sections
        │
        ▼
JVM Bytecode
Example: invokeinterface Lock.lock / Lock.unlock
Role: invoke the ReadLock / WriteLock API
        │
        ▼
Runtime / Language Implementation
Example: ReentrantReadWriteLock.Sync / AbstractQueuedLongSynchronizer
       + Unsafe CAS / LockSupport
Role: implement Reader / Writer state, contention, queuing, and wake-up
        │
        ▼
OS Thread
Example: Runnable / Parked Java Thread
Role: block on contention and become runnable again after unpark
        │
        ▼
x86-64 Hardware
Example: Atomic Instruction / Cache Coherence / Memory Ordering
Role: provide the hardware basis for Atomicity, Visibility, and Ordering

2.2 End-to-End Path from Write Lock to Read Lock

Continue with counter / ready: Thread A writes under the Write Lock, then Thread B reads under the Read Lock.

Mermaid 图表

The sequence diagram keeps only the cross-layer path. The Runtime / Language Implementation layer can be expanded into the Reader / Writer contention flow below:

Mermaid 图表

Two counters are encoded in the synchronization state:

high 32 bits
  → shared count
  → number of Readers

low 32 bits
  → exclusive count
  → Writer reentrancy count

In JDK 25, ReentrantReadWriteLock.Sync extends AbstractQueuedLongSynchronizer. Readers use Shared mode; Writers use Exclusive mode; both ultimately contend on the same 64-bit state.

See OpenJDK 25 ReentrantReadWriteLock.java and AbstractQueuedLongSynchronizer.java.

2.3 Guarantees at the Runtime / Language Implementation Layer

2.3.1 Atomicity

Atomicity comes from the state transitions in the flowchart:

Writer
  → CAS exclusive count
  → exclusive write critical section after success

Reader
  → CAS shared count
  → multiple Readers may succeed independently

A Writer fails if Readers are present or another thread owns the Write Lock.

If the current thread already owns the Write Lock, acquisition is reentrant and increments the exclusive count.

A Reader fails if another thread owns the Write Lock; otherwise it increments the shared count with CAS.

The runtime therefore enforces:

Read + Read
  → multiple Shared acquires allowed

Read + Write
  → conflict

Write + Write
  → conflict

2.3.2 Visibility and Ordering

AbstractQueuedLongSynchronizer stores its state in a volatile long:

private volatile long state;

Its basic operations have these semantics:

getState()
  → volatile read

setState()
  → volatile write

compareAndSetState()
  → volatile read + volatile write

Releasing the Write Lock updates state. A later Reader reads and CASes the same state while acquiring the Read Lock:

Writer writes
    ↓
release Write Lock
    ↓
volatile / CAS state synchronization
    ↓
acquire Read Lock
    ↓
Reader reads

This is where the language-level Visibility and Ordering guarantees are carried by the Runtime / Language Implementation layer.

2.4 Guarantees at the Hardware Layer

At the hardware layer, the implementation reduces to the same three capabilities used throughout the series:

Atomicity
  → Atomic Instruction
  → x86-64: LOCKed CAS and other atomic RMW operations

Visibility
  → Cache Coherence
  → other CPU cores cannot keep using an invalidated old cache line indefinitely

Ordering
  → x86-64 Memory Ordering + LOCKed RMW / required ordering constraints
  → preserve the memory-access order required around lock-state transitions

compareAndSetState() ultimately uses Unsafe.compareAndSetLong, which HotSpot lowers to an atomic CAS on the target architecture.

Read-write locks add no new hardware primitive. They build the rule “multiple Readers or one Writer” on top of atomic state plus thread waiting / wake-up.


3. How Does the Go Runtime Implement sync.RWMutex?

Go’s standard library directly provides sync.RWMutex.

It shares the same core Reader / Writer exclusion model as Java ReadWriteLock: multiple Readers may proceed concurrently, while a Writer must be exclusive. The complete API semantics are not identical.

The most important differences are:

Java ReentrantReadWriteLock Go sync.RWMutex
Read-lock reentrancy Supported Recursive RLock is not supported
Write-lock reentrancy Supported Not supported
Write → Read downgrade Supported Not supported
Read → Write upgrade Not supported Not supported
Fairness Optional fair / nonfair mode No fair mode

The Go Memory Model states:

“For any call to RLock, there exists an n such that the n’th call to Unlock ‘synchronizes before’ that call to RLock.”

3.1 Layers

Go Source Code
Example: rw.RLock() / rw.Lock()
Role: declare read and write critical sections
        │
        ▼
Go Implementation
Example: sync.RWMutex / sync.Mutex / sync/atomic / Runtime Semaphore
Role: implement Reader / Writer state, contention, waiting, and wake-up
        │
        ▼
Goroutine / Scheduler
Example: Runnable / Parked Goroutine
Role: park a contending Goroutine and schedule it again after wake-up
        │
        ▼
x86-64 Hardware
Example: Atomic Instruction / Cache Coherence / Memory Ordering
Role: provide the hardware basis for Atomicity, Visibility, and Ordering

The important distinction from Java is that runtime_SemacquireRWMutex* parks a Goroutine rather than directly blocking one fixed OS thread.

3.2 End-to-End Path from Write Lock to Read Lock

Continue with the same counter / ready example:

var counter int
var ready bool
var rw sync.RWMutex

Goroutine A writes first, then Goroutine B reads:

Mermaid 图表

The sequence diagram keeps only the cross-layer path. The Go Implementation layer can be expanded into the Reader / Writer contention flow below:

Mermaid 图表

The core state is:

w
  → serialize Writers

readerCount
  → >= 0: number of active Readers
  → < 0: a Writer is pending

readerWait
  → number of departing Readers the Writer is still waiting for

readerSem / writerSem
  → Reader / Writer waiting and wake-up

See Go’s sync/rwmutex.go.

3.3 Guarantees at the Runtime / Language Implementation Layer

3.3.1 Atomicity

Atomicity comes from two core state operations in the flowchart.

A Reader enters through:

readerCount.Add(1)

That update must be atomic. Multiple Readers may increment the count independently.

A Writer first serializes against other Writers through the internal w Mutex, then executes:

readerCount.Add(-rwmutexMaxReaders)

to mark a pending Writer.

The runtime therefore enforces:

Read + Read
  → multiple Readers allowed

Read + Write
  → new Readers block once a Writer is pending

Write + Write
  → serialized by w Mutex

Unlike Java, Go does not encode every state in one state word. It combines a Mutex + Atomic Counter + Semaphore to implement the same Reader / Writer rules.

3.3.2 Visibility and Ordering

Go defines RWMutex synchronization at the API level; it should not be inferred from readerCount atomics alone.

The documented relations include:

Unlock
  → synchronized-before
  → the corresponding later RLock

RUnlock
  → synchronized-before
  → the corresponding later Lock

Using the same counter / ready example:

Goroutine A                           Goroutine B

counter = 1
ready = true
    │
    ▼
Unlock() ── synchronized-before ──► RLock() returns
                                        │
                                        ▼
                                    read ready == true
                                    read counter == 1

So once B acquires the Read Lock through this synchronization edge, A’s writes under the Write Lock are visible to B in the required order.

Internally, the Mutex, runtime semaphores, and atomic state implement this path. The readerCount.Add operation alone should not be interpreted as synchronizing Readers with one another.

3.4 Guarantees at the Hardware Layer

At the hardware layer, the implementation reduces to the same three capabilities used throughout the series:

Atomicity
  → Atomic Instruction
  → x86-64: LOCK XADDL / LOCK CMPXCHGL and other atomic RMW operations

Visibility
  → Cache Coherence
  → other CPU cores cannot keep using an invalidated old cache line indefinitely

Ordering
  → x86-64 Memory Ordering + LOCKed RMW
  → provide the ordering constraints required by the synchronization path

readerCount.Add and readerWait.Add ultimately use the x86-64 atomic-add path. The internal w Mutex reuses the CAS / atomic machinery already covered in the Mutex implementation article.

So sync.RWMutex introduces no new hardware primitive. It combines atomics, a Mutex, and runtime semaphores to implement Reader / Writer semantics.


4. CPython: Why Is There No Read-Write Lock in the Standard Library?

Python is capable of implementing a Read-Write Lock. In fact, CPython had an early proposal to add RWLock to threading: Issue 8800.

The proposal never made it into the standard library. Looking at the long-running Issue 8800 discussion, at least two recurring obstacles stand out.

1. Under traditional CPython, the GIL limited the benefit of a Read-Write Lock

A CPython developer summarized the problem directly:

“since the GIL serializes everything anyway, this isn’t likely to benefit many situations”

The basic trade-off was:

Main benefit of a Read-Write Lock
  → multiple Readers run in parallel

Traditional CPython
  → the GIL largely serializes execution of Python code

So even if multiple threads could acquire the Read Lock together, they usually could not execute Python code in parallel.

RWLock was not useless: I/O and C extensions that release the GIL could still benefit. But for general CPython threading, its advantage was much less direct than in Java or Go.

Free-threaded CPython changes this point. Without the GIL, multiple Python threads can truly execute in parallel, so concurrent Readers become substantially more valuable.

That changes the first part of the trade-off. But even if the benefit of RWLock becomes clearer, the standard library still faces the second problem: there was no consensus on the API and scheduling semantics.

2. Issue 8800 did not reach consensus on the API or scheduling policy

The discussion had two main classes of disagreement.

The first was the API shape:

one unified RWLock
or
two associated Shared / Exclusive Lock objects

One proposal argued that:

“having two different lock primitives … would be much more flexible, pythonic”

Others preferred a single RWLock abstraction.

The second was the scheduling policy:

Reader priority or Writer priority?
Fair / FIFO?
May new Readers enter while a Writer waits?
Can a Read Lock be upgraded?
Can a Write Lock be downgraded?

Those choices directly affect RWLock behavior and API semantics, and Issue 8800 never converged on one standard design.

So the obstacle was not implementation feasibility. There was no consensus on which RWLock API and policy the standard library should permanently define.

The current threading standard library still provides Lock, RLock, Condition, Semaphore, Event, and related primitives, but no Read-Write Lock.

Applications that need RWLock semantics generally have two options:

Use a third-party / custom RWLock
        │
        └── obtain explicit Read Lock / Write Lock semantics

Build on threading.Lock + Condition
        │
        └── define Reader / Writer counts and scheduling policy yourself

The original Issue 8800 proposal itself implemented RWLock by composing a condition variable and a lock.

What Python lacks is therefore a standardized RWLock API and policy in the standard library, not the ability to implement Read-Write Lock semantics.


5. Common Pattern Across the Two Implementations

Strip away the APIs and compare what Java and Go must solve.

5.1 Who Gets In? — Atomically Maintaining Reader / Writer State

A Mutex only needs to answer whether the lock is held.

A ReadWriteLock must maintain:

Reader Count
+
Writer State

Java:

64-bit state
  → shared count
  → exclusive count

Go:

readerCount
+
Writer Mutex
+
readerWait

The encoding differs, but the admission rules are the same:

Reader
  → no conflicting Writer

Writer
  → no other Writer
  → no Reader

Those state transitions must be atomic. Otherwise competing execution units could observe stale state and enter incorrectly.

5.2 What Happens on Conflict? — Park / Wakeup

After a state check fails, a waiter cannot spin forever.

Both implementations reduce to:

check Reader / Writer state
        │
        ├── allowed
        │      └── update state and enter
        │
        └── conflict
               │
               ▼
           queue / wait
               │
               ▼
              Park
               │
               ▼
       relevant holder releases
               │
               ▼
            Wakeup
               │
               └── Retry

The parked unit differs:

Java
  → AQLS Sync Queue
  → LockSupport.park()
  → OS Thread Parked

Go
  → Runtime Semaphore
  → Park Goroutine
  → Scheduler runs another Goroutine

So the common pattern is Conflict → Park → Wakeup → Retry. Java parks a Thread; Go parks a Goroutine.

5.3 Why Can a Later Reader See an Earlier Writer’s Data?

The final problem is the synchronization boundary:

Writer writes
    │
    ▼
Write Unlock [Release]
    │
    ▼
Read Lock [Acquire]
    │
    ▼
Reader reads

Java implements this boundary through AQLS volatile / CAS state transitions.

Go defines the semantics through RWMutex synchronized-before rules and implements them with its internal Mutex, runtime semaphores, and atomic operations.

Both implementations must guarantee the same result:

Writes before the Writer releases the lock precede reads performed after the later Reader acquires it.


6. How ReadWriteLock and Mutex Relate

Both protect critical sections. The difference is whether Readers and Writers are distinguished:

Mutex ReadWriteLock
Core problem mutual exclusion for a critical section distinguish read and write critical sections
Atomicity entire critical section is exclusive conflicting operations are exclusive; Readers may coexist
Visibility Yes Yes
Ordering Yes Yes
Read–Read Exclusive Concurrent
Read–Write Exclusive Exclusive
Write–Write Exclusive Exclusive
Typical use general shared-state protection read-heavy workloads with real contention in read-side critical sections

A compact model is:

Mutex
  → everyone uses one Exclusive entry

ReadWriteLock
  → Readers use the Shared entry
  → Writers use the Exclusive entry
  → only Read + Read is relaxed

ReadWriteLock is not a “stronger” Mutex.

It refines the exclusion boundary so non-conflicting Readers may proceed concurrently.


7. Next: AQS / AQLS — Java’s Unified Synchronizer Framework

This article has already crossed the same path several times:

ReentrantReadWriteLock
        ↓
Sync
        ↓
AbstractQueuedLongSynchronizer

That is not merely an implementation detail of one lock. It belongs to the lower-level synchronizer framework used throughout Java’s concurrency library.

The next article goes deeper into the AQS family:

AQS / AQLS
  → state
  → Exclusive / Shared
  → CLH-style Queue
  → acquire / release
  → park / unpark
  → Condition Queue

In JDK 25, ReentrantReadWriteLock specifically uses AbstractQueuedLongSynchronizer. Its core model is the same as AbstractQueuedSynchronizer, but the synchronization state is a long.

Go has no single framework directly equivalent to AQS.

Instead, Go synchronization primitives tend to compose lower-level mechanisms directly:

Atomic State
+
Mutex
+
Runtime Semaphore
+
Goroutine Scheduler

For example, sync.RWMutex directly combines w, readerCount, readerWait, readerSem, and writerSem rather than inheriting from a generic AQS-style synchronizer.

AQS / AQLS is therefore a major layer in Java’s lock implementation stack and deserves its own article next.

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub