Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

6.2 · Mutex, RwLock, and Poison Recovery

Domain 6 — Concurrency and Parallelism Duration: ~15 minutes Library components: std::sync::Mutex, std::sync::MutexGuard, std::sync::RwLock, std::sync::RwLockReadGuard, std::sync::RwLockWriteGuard, std::sync::Condvar, std::sync::PoisonError, std::sync::TryLockError

Introduction

In most languages, only a convention connects a mutex to the data that it protects. If you forget to lock the mutex, you have a data race. The Mutex<T> of Rust owns the data. The only path to the T is through lock(). Thus access without a lock is a compile error, not a failure at runtime.

This tutorial shows:

  • Mutex<T>: lock, try_lock, the MutexGuard RAII pattern, and into_inner.
  • RwLock<T>: many concurrent readers or one writer, and when that trade-off is an advantage.
  • Poisoning: what it means, why lock() returns a Result, and how to recover with PoisonError::into_inner and clear_poison.
  • Condvar: how to block until a condition is true, without a busy-wait loop.

Mutex<T>: The Lock Owns the Data

Figure: Mutex Lock Lifecycle

The usual example is a counter that several threads share. This version is safe:

// The mutex owns the u64. `Arc` gives each thread shared ownership of the mutex.
let counter = Arc::new(Mutex::new(0_u64));

let mut handles = Vec::new();
for _ in 0..4 {
    let counter = Arc::clone(&counter);   // one Arc handle for each thread
    handles.push(thread::spawn(move || {
        for _ in 0..1_000 {
            // `lock` blocks until the mutex is free. The guard derefs to &mut u64.
            let mut guard = counter.lock().unwrap();
            *guard += 1;
        }  // the guard drops here and unlocks the mutex, even on panic or early return
    }));
}
for handle in handles {
    handle.join().expect("worker should not panic");
}

assert_eq!(*counter.lock().unwrap(), 4_000);   // deterministic after the joins

Three points are important:

  • Arc<Mutex<T>> is the standard pair. Arc shares ownership across threads, and Mutex serializes access. Tutorial 6.8 explains why you need the two types.
  • MutexGuard is RAII. The unlock is in Drop, so no code path can omit it. To release the lock early, before slow work that does not need the lock, call drop(guard) explicitly.
  • lock() returns Result because of poisoning (slide 5). .unwrap() is the idiomatic default, not a shortcut: the program refuses to run on data that is possibly corrupt.

When the sharing ends, you can recover the data without a lock. Arc::try_unwrap(arc) returns the Mutex when only one owner remains. mutex.into_inner() returns the T. Exclusive ownership proves that no other thread can race with you.

try_lock: Refusing to Wait

try_lock() never blocks. It immediately returns Ok(guard) or Err(TryLockError::WouldBlock). Use it for opportunistic work. For example: flush the cache if no thread uses it, or else skip this cycle.

let busy = Mutex::new(-1_i32);
let held = busy.lock().unwrap();   // this thread holds the lock

// The mutex is not re-entrant: a second attempt from the same thread fails.
match busy.try_lock() {
    Err(TryLockError::WouldBlock) => { /* skip, try again in the next cycle */ }
    Err(TryLockError::Poisoned(_)) => unreachable!("nothing has panicked"),
    Ok(_) => unreachable!("mutex is already held"),
}
drop(held);                         // release the lock
assert!(busy.try_lock().is_ok());   // now `try_lock` succeeds immediately

The Mutex of std is not re-entrant. A second lock from the same thread never returns: the documentation says that it might panic or deadlock. A second try_lock fails. Thus the example above is deterministic without a second thread. The same property is a risk in real programs. Do not call a function that locks a mutex from inside a critical section that holds the same lock.

RwLock<T>: Many Readers or One Writer

RwLock divides access into read() and write(). read() is shared: any number of read guards can exist concurrently. write() is exclusive. RwLock is a good selection for read-mostly data. An example is a service configuration that each request reads and only a redeploy rewrites.

let telemetry = RwLock::new(vec![10_u32, 20, 30]);

// Two read guards exist at the same time, in the same thread.
let reader_a = telemetry.read().unwrap();
let reader_b = telemetry.read().unwrap();
assert_eq!(reader_a.iter().sum::<u32>(), 60);
assert_eq!(reader_b.len(), 3);

// While a reader exists, a writer cannot enter. `try_write` fails immediately.
assert!(matches!(telemetry.try_write(), Err(TryLockError::WouldBlock)));

drop(reader_a);
drop(reader_b);
// No readers remain, so the writer gets exclusive access.
telemetry.write().unwrap().push(40);
assert_eq!(*telemetry.read().unwrap(), [10, 20, 30, 40]);

// `06_06_rwlock.rs` then shares an Arc<RwLock<Config>> between 4 reader threads
// and 1 writer thread. The readers never block each other.

The reader/writer rule is the rule of the borrow checker (&T xor &mut T), which RwLock applies at runtime. RwLock is the thread-safe equivalent of RefCell: it blocks where RefCell panics. Tutorial 7.3 describes the remainder of that family.

When RwLock is not better than Mutex: If writes are frequent, or critical sections are very small, the extra bookkeeping costs more than it saves. This bookkeeping includes the reader count and writer preference. Std also makes no fairness guarantee. It inherits the policy of the OS lock, so a continuous flow of readers can starve writers on some platforms. Use Mutex by default. Use RwLock when a profile shows that reads are much more frequent than writes and the read sections do real work.

Poisoning: The Signal of a Panic

A mutex becomes poisoned when a thread panics while it holds the guard. The reason is that the panic may have interrupted a multi-step update, which leaves the protected data with a broken invariant. Poisoning makes that suspicion visible: each subsequent lock() returns Err(PoisonError).

Figure: Poison Recovery Flow

Poisoning is advisory, not destructive. The PoisonError contains the guard that you requested:

// ledger: Arc<Mutex<Ledger>>, initially Ledger { enqueued: 10, processed: 8 }.
// A worker added 5 to `enqueued`, then panicked while it held the guard.
let mut guard = match ledger.lock() {
    Ok(guard) => guard,
    Err(poisoned) => poisoned.into_inner(),   // get the guard from the PoisonError
};
guard.enqueued -= 5;             // repair the half-applied update: 15 becomes 10 again
drop(guard);

assert!(ledger.is_poisoned());   // the recovery does NOT clear the flag
ledger.clear_poison();           // this call clears the flag
assert!(ledger.lock().is_ok());

RwLock becomes poisoned in the same way, with one difference: only a writer that panics poisons it. A reader that panics could not corrupt the data, so it does not poison the lock.

06_07_poison_recovery.rs prints:

mutex poisoned after worker panic: true
recovering: taking the guard out of the PoisonError
state at recovery: enqueued=15 processed=8
after clear_poison: enqueued=10 processed=8

All assertions passed.

Condvar: Sleeping Until a Condition Holds

A Condvar works together with a Mutex that protects the actual condition state. A waiter atomically releases the lock and sleeps. A notifier changes the state and wakes the waiters. Use wait_while: it checks the predicate again on each wakeup. Thus spurious wakeups (the OS can produce them) are harmless:

// jobs: Mutex<VecDeque<u32>>, ready: Condvar. The two threads share the pair in an Arc.

// Consumer: sleep while the queue is empty.
let mut guard = jobs.lock().unwrap();
// `wait_while` releases the lock while it sleeps. It returns the guard when the
// predicate is false.
guard = ready.wait_while(guard, |q| q.is_empty()).unwrap();
let job = guard.pop_front().expect("predicate guarantees non-empty");
drop(guard);   // release the lock BEFORE the slow work

// Producer: change the state, THEN notify. Never use the opposite order.
// job: u32 — the number of the next job
jobs.lock().unwrap().push_back(job);
ready.notify_one();
  • notify_one wakes one waiter (work queues). notify_all wakes all the waiters (startup gates, shutdown broadcasts).
  • wait_timeout limits the sleep. It returns the guard and a WaitTimeoutResult, and the timed_out() method tells you why the thread woke.
  • The order "change the state, then notify" is important. If you notify first, the waiter may check the predicate again, find the old state, and sleep. Then it misses the only wakeup that it could get.

Condvar is the most general tool for a wait. Channels (Tutorial 6.5) and Barrier (Tutorial 6.7) package the same mechanism for their specific cases, and they are easier to use.

Choosing: Mutex, RwLock, or Something Else

SituationUse this
Any shared mutable state (the default selection)Mutex<T>
Read-mostly data, where the read sections do real workRwLock<T>
One value, one operation (counter, flag)Atomic* (Tutorial 6.3)
Data that flows in one direction between threadschannels (Tutorial 6.5)
A wait on an arbitrary predicateCondvar
One-time initializationOnceLock/LazyLock (Tutorial 6.6)

Two habits prevent most lock bugs:

  • Keep critical sections small: compute outside, change the data inside.
  • Never call unknown code (callbacks, Display implementations, loggers that allocate) while you hold a guard.

Summary

ConceptKey point
Mutex<T> owns its dataAccess is only through lock(). Access without a lock does not compile.
MutexGuardRAII: Drop unlocks the mutex on each code path.
lock().unwrap()Idiomatic: it propagates poisoning and does not run on corrupt data.
try_lockDoes not block. Returns WouldBlock when a thread holds the lock. Use it for opportunistic work.
Not re-entrantA second lock from the same thread never returns. Do not lock the same mutex again in a critical section.
RwLockConcurrent readers xor one writer. Best for read-mostly data.
Writer starvationNo fairness guarantee. The behavior depends on the platform.
PoisoningA panic while a thread holds the guard causes it. It is an advisory flag, not data loss.
PoisonError::into_innerReturns the guard. Repair the invariants, then call clear_poison().
Condvar::wait_whileChecks the predicate again on each wakeup, so spurious wakeups are harmless.
Change the state, then notifyThe one ordering rule of Condvar protocols.

Code Examples

FileDescription
06_05_mutex_basics.rslock, guard RAII, try_lock/WouldBlock, into_inner
06_06_rwlock.rsConcurrent readers, try_write, read-mostly config workload
06_07_poison_recovery.rsDeliberate poisoning, PoisonError::into_inner, clear_poison
06_08_condvar.rswait_while queue, notify_all gate, wait_timeout