Concurrency Programming (6): How Atomic Operations Are Implemented — From Runtime to CPU
Traces the implementation paths of Java AtomicInteger, Go sync/atomic, and CPython's internal atomic operations to show how atomic RMW, CAS, and memory ordering reach the CPU.

Table of Contents
- 0. What Does This Article Answer Next?
- 1. How Does HotSpot Implement AtomicInteger?
- 2. How Is sync/atomic Implemented in Go?
- 3. How Does CPython Implement Internal Atomics?
- 4. The Common Pattern Across the Three Implementations
- 5. Next: volatile
0. What Does This Article Answer Next?
The previous article explained the Atomicity, Visibility, and Ordering guarantees provided by atomic APIs at the language level. This article follows those guarantees down through the runtime and compiler to the CPU.
We continue to use the same two examples:
counter++
and:
counter = 1
ready = true
The implementation path can be summarized as:
Language API and memory model
↓
Compiler / Runtime
↓
CPU Atomic Instruction / Memory Ordering
1. How Does HotSpot Implement AtomicInteger?
1.1 Layers
Java Source Code
Example: AtomicInteger.incrementAndGet() / compareAndSet() / get()
Role: express atomic operations
│
▼
JVM Bytecode
Example: invokevirtual AtomicInteger.incrementAndGet() / compareAndSet() / get()
Role: invoke AtomicInteger methods; there is no dedicated bytecode for atomic operations
│
▼
JVM Implementation (HotSpot)
Example: Intrinsic / C1 / C2
Role: lower atomic operations to the target architecture
│
▼
x86-64 Hardware
Example: Atomic Instruction / Cache Coherence / Fence
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
1.2 End-to-End Implementation Path for AtomicInteger
The main atomic operations can be expanded as follows:
1.3 Guarantees at the Runtime / Language Implementation Layer
1.3.1 Atomicity
At the Runtime / Language Implementation layer, Atomicity is preserved through two kinds of indivisible operations:
Atomic Add
→ Unsafe.getAndAddInt
→ HotSpot Intrinsic / C1 / C2
→ Atomic Get-And-Add
CAS
→ Unsafe.compareAndSetInt
→ HotSpot Intrinsic / C1 / C2
→ Atomic CAS
HotSpot preserves their atomic semantics through compilation and lowering; the target architecture then provides the concrete machine instructions.
1.3.2 Visibility and Ordering
Java atomic methods carry explicit memory effects in addition to indivisible RMW semantics. incrementAndGet() / getAndAdd() use the volatile read/write effects of VarHandle.getAndAdd, while get() has volatile-read semantics.
HotSpot must preserve those memory-ordering constraints through compilation and optimization.
1.4 Guarantees at the Hardware Layer
At the hardware layer, x86-64 memory ordering, cache coherence, and locked RMW operations provide the required guarantees. Typical mappings include LOCK XADD for atomic add and LOCK CMPXCHG for CAS.
Atomicity
→ Atomic Instruction
→ HotSpot / x86-64: LOCK XADD / LOCK CMPXCHG
Visibility
→ Cache Coherence
→ Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)
Ordering
→ Compiler Ordering + Hardware Memory Ordering
→ HotSpot Memory Effects + x86-64 Locked RMW / Load rules
LOCK XADD and LOCK CMPXCHG are typical x86-64 mappings, not a guarantee that every JVM version or architecture emits the same instruction sequence.
See OpenJDK AtomicInteger.java, Unsafe.java, and HotSpot atomicAccess_linux_x86.hpp.
2. How Is sync/atomic Implemented in Go?
2.1 Layers
Go Source Code
Example: atomic.Int64.Add() / CompareAndSwap() / Load()
Role: express atomic operations
│
▼
Go Implementation
Example: sync/atomic + internal/runtime/atomic
Role: implement the atomic API and target-specific atomic primitives
│
▼
x86-64 Hardware
Example: Atomic Instruction / Cache Coherence / Fence
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
2.2 End-to-End Implementation Path for atomic.Int64
The main atomic operations can be expanded as follows:
2.3 Guarantees at the Runtime / Language Implementation Layer
2.3.1 Atomicity
At the Runtime / Language Implementation layer, Atomicity is represented by Go’s atomic primitives:
Atomic Add
→ sync/atomic
→ internal/runtime/atomic.Xadd64
CAS
→ sync/atomic
→ internal/runtime/atomic.Cas64
These operations already carry indivisible atomic semantics at this layer; the target architecture provides the concrete machine instructions.
2.3.2 Visibility and Ordering
The Go Memory Model defines synchronized-before for atomic operations and requires all atomic operations to behave as though they execute in some sequentially consistent order. The compiler and Runtime must preserve that observable behavior through lowering and optimization.
2.4 Guarantees at the Hardware Layer
At the hardware layer, amd64 maps those primitives to instructions such as LOCK XADDQ and LOCK CMPXCHGQ, while loads follow the architecture’s ordering rules.
Atomicity
→ Atomic Instruction
→ Go / amd64: LOCK XADDQ / LOCK CMPXCHGQ
Visibility
→ Cache Coherence
→ Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)
Ordering
→ Compiler Ordering + Hardware Memory Ordering
→ Go Atomic semantics + amd64 Locked RMW / Load rules
See sync/atomic, internal/runtime/atomic, and atomic_amd64.s.
3. How Does CPython Implement Internal Atomics?
This section covers only the atomic operations used internally by the current CPython runtime. They are not a public application-level Python API.
3.1 Layers
CPython Runtime Source Code
Example: internal shared-state updates
Role: call internal atomic operations
│
▼
CPython Implementation
Example: _Py_atomic_* + GCC / Clang __atomic_*
Role: implement atomic operations and memory-order semantics, then lower them to the target architecture
│
▼
x86-64 Hardware
Example: Atomic Instruction / Cache Coherence / Fence
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
3.2 End-to-End Implementation Path for CPython Atomics
The main internal atomic operations can be expanded as follows:
3.3 Guarantees at the Runtime / Language Implementation Layer
3.3.1 Atomicity
At the Runtime / Language Implementation layer, CPython expresses Atomicity through its internal atomic abstraction and the compiler’s atomic builtins:
Atomic Add
→ _Py_atomic_add_*
→ __atomic_fetch_add
CAS
→ _Py_atomic_compare_exchange_*
→ __atomic_compare_exchange_n
These operations carry explicit atomic semantics before the compiler selects the final instruction sequence for the target architecture.
3.3.2 Visibility and Ordering
CPython’s _Py_atomic_* layer supports multiple memory orders. Current implementations use sequentially consistent operations by default, alongside variants such as _relaxed, _acquire, and _release.
Visibility and Ordering therefore depend on the memory order selected at the call site:
CPython Runtime
→ choose _Py_atomic_* Memory Order
→ GCC / Clang __atomic_* order
3.4 Guarantees at the Hardware Layer
At the hardware layer, the compiler maps those atomic operations and memory-order constraints to x86-64 loads, stores, locked RMW instructions, and the architecture’s ordering rules.
Mapped back to the hardware model:
Atomicity
→ Atomic Instruction
→ x86-64: LOCK XADD / LOCK CMPXCHG
Visibility
→ Cache Coherence
→ Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)
Ordering
→ Compiler Atomic Memory Order + Hardware Memory Ordering
→ __atomic_* order + x86-64 ordering rules
This is an implementation property of the CPython runtime, not a Python language memory model.
See pyatomic.h, pyatomic_gcc.h, and a real use site in Python/lock.c.
4. The Common Pattern Across the Three Implementations
Java, Go, and CPython expose different APIs, but their atomic implementations converge on the same three problems:
1. How can one update become indivisible?
2. If a conditional update fails, who decides whether to retry?
3. How do ordinary reads and writes around the atomic operation gain the required visibility and ordering?
4.1 How Does a Single-Variable Update Become Indivisible? — Atomic RMW
An ordinary read → modify → write consists of several steps that can interleave with other execution units. Atomic Add, Exchange, CAS, and other RMW operations collapse that shared-state update into a single atomic CPU operation:
Language / Runtime Atomic
↓
Compiler / Intrinsic
↓
Atomic RMW Instruction
↓
Cache Coherence
On x86-64, typical examples are LOCK XADD and LOCK CMPXCHG.
4.2 What Happens When a Conditional Update Fails? — CAS / Retry
CAS performs one atomic compare-and-conditional-write operation:
expected still matches
→ update succeeds
expected is stale
→ update fails
CAS itself does not imply retry. compareAndSet() / CompareAndSwap() may simply return failure. updateAndGet(), some Runtime algorithms, or caller-written CAS loops are the layers that decide to reload, recompute, and retry.
CAS
= one atomic conditional update
CAS Loop
= CAS + algorithm-level retry after failure
4.3 Why Can Atomics Also Establish Visibility and Ordering?
An indivisible RMW operation solves Atomicity, but not the whole memory-ordering problem.
Java VarHandle memory effects, the Go Memory Model’s atomic rules, and the memory order selected by CPython _Py_atomic_* also constrain:
Compiler Reordering
+
CPU Memory Ordering
+
Cache Coherence
So the full path remains:
Language / Runtime defines Memory Semantics
↓
Compiler preserves ordering constraints
↓
CPU uses Atomic Instruction / Load / Store / Fence mechanisms
↓
Cache Coherence propagates shared state
↓
Program receives the corresponding Visibility and Ordering
This is also where atomics and the previous Mutex article converge at the hardware layer: both rely on Atomic Instructions, Cache Coherence, and Memory Ordering. A Mutex additionally needs waiting and wake-up paths under sustained contention; an atomic operation itself has no Park / Wakeup path.
4.4 How Do Atomics Differ from Mutexes?
The simplest distinction is:
An atomic operation makes one shared-state operation indivisible; a Mutex protects a block of code.
4.4.1 One State Update, Two Approaches
Suppose all we need is to increment one counter.
With a Mutex:
Lock
↓
counter++
↓
Unlock
With an Atomic:
Atomic Add(counter, 1)
The Mutex approach is:
let one execution unit enter
↓
run the protected code
↓
leave the critical section
The Atomic approach is:
make this update itself
indivisible
If the problem is truly one Add, CAS, Swap, Load, or Store, there is no need to wrap that atomic operation in an additional critical section.
4.4.2 Why Can’t Multiple Atomics Replace One Mutex?
Suppose one logical operation must update two pieces of state:
balance -= 100
count++
Making both variables atomic gives us:
Atomic balance -= 100
↓
← another thread may observe this intermediate state
↓
Atomic count++
Each update is indivisible on its own, but the two updates are not atomic as a group.
A Mutex can protect both updates as one critical section:
Lock
↓
balance -= 100
count++
↓
Unlock
Until that critical section completes, another contender cannot enter the same protected region.
So once the requirement changes from “update one state” to “these steps must complete together,” a Mutex usually expresses the intent more directly.
4.4.3 Which Should You Use?
Look at what needs protection:
one state operation
Add / CAS / Swap / Load / Store
↓
Atomic
a block of logic
read → decide → update multiple states
↓
Mutex
The two are not independent. A Mutex commonly uses atomic operations or CAS internally to decide who acquires the lock:
Atomic / CAS
↓
compete for lock state
↓
Mutex
↓
protect a critical section
So the relationship can be summarized as:
Atomic: one shared-state operation
Mutex: a critical section built on lower-level primitives such as atomic operations / CAS
5. Next: volatile
Atomic operations solve the atomicity problem for RMW updates such as counter++. If no atomic update is needed and the goal is only to use ready to publish an already-written counter, Java provides volatile. The next article explains how it provides visibility and ordering, while also explaining why Go and Python do not have a corresponding volatile keyword.
Discussion
Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub