V. Coordination & Consensus
Coordination is what you reach for when independent nodes must agree — on who is the leader, on whether a transaction committed, on who holds a lock. It is also the most expensive thing you can do in a distributed system, because agreement requires rounds of communication that fail exactly when you need them most. This part covers distributed transactions and their escape hatches (2PC, sagas), the consensus algorithms that make agreement possible at all (Paxos, Raft), and why distributed locks are a correctness trap without fencing.
Topics Covered
Section titled “Topics Covered”- 5.1. Distributed Transactions: Rebuilding atomicity across services: two-phase commit, the saga pattern, compensating transactions, and the outbox pattern.
- 5.2. Consensus Algorithms: Agreeing on a value despite failures: the FLP impossibility, Paxos, Raft, Zab, etcd, and Byzantine fault tolerance.
- 5.3. Distributed Locking: Mutual exclusion across machines: lease-based locking, fencing tokens, and the Redlock correctness debate.