Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

6.4 · Memory Ordering: Relaxed, Acquire/Release, and SeqCst

Domain 6 — Concurrency and Parallelism Duration: ~15 minutes Library components: std::sync::atomic::Ordering, std::sync::atomic::fence, std::sync::atomic::compiler_fence

Introduction

Compilers reorder instructions, and CPUs reorder memory operations. They do this aggressively and invisibly, and the result is correct for single-threaded code. When a second thread observes your memory, the order of events is no longer clear. Each atomic operation takes an Ordering argument. With this argument, you request exactly as much cross-thread ordering as you need.

This tutorial shows:

  • The guarantee of each Ordering: Relaxed, Acquire, Release, AcqRel, SeqCst.
  • The happens-before relation, with the message-passing pattern (a payload and a ready flag) as the example.
  • The difference between compare_exchange and compare_exchange_weak, and the retry loop.
  • fence and compiler_fence.
  • The common bugs: the assumption that Relaxed orders unrelated memory, and orderings that do not match between the store side and the load side.

What Each Ordering Guarantees

Each atomic operation is itself indivisible with every ordering. The orderings differ in the constraint that the operation puts on the other memory accesses near it:

OrderingForGuarantee
Relaxedloads and storesAtomicity only. No cross-thread ordering of other memory.
ReleasestoresThe compiler and the CPU cannot move a write from before this store to after it.
AcquireloadsThe compiler and the CPU cannot move a read or a write from after this load to before it.
AcqRelRMW (read-modify-write) operationsThe two halves: Acquire for the read and Release for the write.
SeqCstall operationsAcquire/Release, plus one global order of all SeqCst operations.

The pairing rule: a Release store synchronizes-with an Acquire load of the same atomic that observes the stored value. This edge and the program order in each thread together make the happens-before relation. The happens-before relation decides what a thread is guaranteed to see.

Relaxed is correct for independent counters (Tutorial 6.3), because a counter needs only its own consistency. When a flag publishes other memory, Relaxed on that flag is a bug.

The Message-Passing Pattern

The producer thread writes data and then sets a flag. The consumer thread waits for the flag and then reads the data. Channels, OnceLock, and every hand-off protocol use this pattern.

Figure: Release/Acquire Happens-Before Edge

static PAYLOAD: AtomicU64 = AtomicU64::new(0);       // the data
static READY: AtomicBool = AtomicBool::new(false);   // the flag that publishes the data

// Producer thread:
PAYLOAD.store(42, Ordering::Relaxed);     // (1) write the data
READY.store(true, Ordering::Release);     // (2) then publish it

// Consumer thread:
while !READY.load(Ordering::Acquire) {    // (3) spin until the flag is true
    hint::spin_loop();                    // tells the CPU that this is a spin-wait
}
assert_eq!(PAYLOAD.load(Ordering::Relaxed), 42);   // (4) always sees the 42 from (1)

The chain (1) → (2) → (3) → (4) has no gap:

  • Release keeps (1) before (2).
  • The synchronizes-with edge connects (2) to the load (3) that reads true.
  • Acquire keeps (4) after (3).

The payload accesses can be Relaxed, because the flag supplies the ordering.

The example binary shows two engineering practices:

  • Bounded spinning: spin for a short time with hint::spin_loop(), which emits the pause instruction of the CPU. Then use thread::yield_now(), so that the producer is sure to get scheduler time, even on one core. Never spin at full speed without a limit.
  • Correct by construction, not by chance: on x86, a Relaxed flag would usually pass, because the hardware is strongly ordered. On ARM and POWER, a Relaxed flag can fail in practice. The example runs 200 rounds, but its purpose is not to test the race. The rounds use a guarantee that holds in each round and on each architecture.

06_12_message_passing.rs prints:

message passing: 200 rounds, payload visible every time

All assertions passed.

compare_exchange: Ordering-Aware CAS

compare_exchange(expected, new, success, failure) is a compare-and-swap (CAS) operation. It atomically replaces the value only if the value equals expected. On success, it returns Ok(previous). On failure, it returns Err(actual) and writes nothing.

It takes two orderings: one for success and one for failure. A failed CAS does no store, so the failure ordering cannot be Release or AcqRel.

A one-shot claim, for example leader election:

// leader: AtomicU32, starts as IDLE (0). my_id: u32, the id of this thread (not 0).
// Success ordering: AcqRel. Failure ordering: Acquire.
match leader.compare_exchange(IDLE, my_id, Ordering::AcqRel, Ordering::Acquire) {
    Ok(_) => { /* exactly one thread gets Ok: it is the leader */ }
    Err(current_leader) => { /* each other thread gets the id of the leader */ }
}

The weak retry loop

compare_exchange_weak can fail spuriously: it can report a failure even when the value matched. On LL/SC architectures (ARM, RISC-V), the strong version hides spurious failures with its own internal loop. That internal loop is unnecessary work if your code already retries. The standard pattern is a retry loop around the weak version. This example implements fetch_mul, an operation that the standard library does not have:

// Multiplies the atomic by `factor` and returns the previous value.
fn fetch_mul(atom: &AtomicU32, factor: u32) -> u32 {
    let mut current = atom.load(Ordering::Relaxed);   // the first value to try
    loop {
        let new = current * factor;
        match atom.compare_exchange_weak(current, new, Ordering::Relaxed, Ordering::Relaxed) {
            Ok(previous) => return previous,   // the CAS stored `new`
            // Err holds the value that is in the atomic now. Try again with it.
            Err(actual) => current = actual,
        }
    }
}

General rule: in a loop, use the weak version. For one attempt, use the strong version. The update and try_update methods from Tutorial 6.3 contain this same loop. So does fetch_update, which Rust 1.99 deprecates in favor of try_update.

06_13_compare_exchange.rs prints:

leader election: winner id=1 (varies), winners=1 losers=3
fetch_mul x4 threads: 3 * 2^4 = 48

All assertions passed.

fence and compiler_fence

atomic::fence(ordering) separates the ordering from a specific atomic operation. The producer calls a Release fence and then does any store. The consumer does a load that observes this store and then calls an Acquire fence. This sequence creates the same synchronizes-with edge. The result is one ordering point for a full batch:

// BATCH: [AtomicU64; 4] and PUBLISHED: AtomicBool are statics. values is [11, 22, 33, 44].
// Producer: N Relaxed stores, ONE Release fence, then a Relaxed flag store.
for (slot, value) in BATCH.iter().zip(values) {
    slot.store(value, Ordering::Relaxed);
}
fence(Ordering::Release);                   // orders all the stores above before the flag store
PUBLISHED.store(true, Ordering::Relaxed);   // the publication point

// Consumer: a Relaxed flag loop, then ONE Acquire fence.
while !PUBLISHED.load(Ordering::Relaxed) { hint::spin_loop(); }
fence(Ordering::Acquire);
// All the BATCH slots are now visible: [11, 22, 33, 44].

A fence has a cost one time for each batch. An ordering on each operation has a cost for each operation. Most code never needs a fence. Use a fence when a profile shows many stores before one publication point.

compiler_fence is a different type of tool. It emits no CPU instruction. It only makes sure that the compiler does not reorder memory accesses across it. Its use is ordering on the same core: signal handlers, interrupt handlers, and memory-mapped I/O. It is never a replacement for fence in cross-thread code.

06_14_fences.rs prints:

batch after publication: [11, 22, 33, 44]

All assertions passed.

SeqCst, and the Bugs to Avoid

Acquire/Release creates edges between pairs of operations. SeqCst also puts each SeqCst operation into one total order that all threads agree on. The standard case that needs SeqCst is store buffering (Dekker's algorithm). Two threads each store their own flag and then load the flag of the other thread. With only Acquire/Release, the two threads can both load false. With SeqCst, at least one thread must see true.

If your protocol depends on more than one independent atomic, start with SeqCst.

Figure: Choosing an Ordering

The common bugs:

  1. Relaxed used to order unrelated memory. Relaxed gives a guarantee for the atomic itself and for nothing near it. A payload behind a Relaxed flag is broken in practice on weakly-ordered CPUs. On x86, the bug is not visible until you run the program on ARM.
  2. Mismatched sides. A Release store synchronizes only with an Acquire load of the same atomic. Release on one side and Relaxed on the other side create no edge. Examine the two sides together.
  3. Ordering used to excuse a data race. No ordering makes a non-atomic data race defined behavior. Fences order atomics. They do not make a race on a static mut sound.

The practice: use the weakest ordering that you can justify in a comment. Use SeqCst when the justification is not precise. Make the code correct first. Then weaken the ordering only with evidence.

Summary

ConceptKey point
RelaxedAtomicity only. Correct for independent counters.
Release storeEarlier writes cannot move after it. It is the publish side.
Acquire loadLater accesses cannot move before it. It is the consume side.
synchronizes-withA Release store and an Acquire load of the same atomic that observes the value
happens-beforeProgram order and synchronizes-with edges, combined
AcqRelFor RMW operations that read and also publish
SeqCstAdds one global total order. For protocols with more than one atomic. The safe default.
compare_exchangeTwo orderings. The failure ordering cannot be Release or AcqRel.
compare_exchange_weakIt can fail spuriously. Use it in retry loops.
fenceOrdering for a batch, separate from a single operation
compiler_fenceCompiler only. For signal handlers and MMIO, not for cross-thread synchronization.

Code Examples

FileDescription
06_12_message_passing.rsRelease/Acquire hand-off, bounded spin loop, 200 verified rounds
06_13_compare_exchange.rsOne-shot CAS claim, weak-CAS retry loop (fetch_mul)
06_14_fences.rsBatch publication with fence, compiler_fence, SeqCst notes