13.6. Anatomy of Famous Outages
Every large-scale outage is a distributed-systems lesson written in production. This chapter dissects four landmark incidents — GitHub’s 2018 split-brain, Roblox’s 73-hour Consul cascade, Cloudflare’s 2019 regex meltdown, and Meta’s 2021 BGP self-lockout — reading each from its official postmortem. The analysis is blameless and systemic: the root cause is never an individual but a chain of design decisions, and each incident closes with a defense-in-depth mapping back to the resilience and observability chapters that would have contained it.
Topics Covered
Section titled “Topics Covered”- 13.6.1. GitHub 2018: 43 Seconds of Partition, 24 Hours of Recovery: Analyzes GitHub’s 2018 outage where a 43-second partition triggered 24 hours of data-reconciliation recovery.
- 13.6.2. Roblox 2021: The 73-Hour Consul Outage: Dissects Roblox’s 73-hour outage caused by a Consul cascade under a novel load pattern.
- 13.6.3. Cloudflare 2019: The Regex That Stalled the Edge: Analyzes Cloudflare’s 2019 outage where a catastrophically backtracking regex stalled the entire edge.
- 13.6.4. Meta 2021: BGP, DNS, and Control Plane Centralization: Dissects Meta’s 2021 outage where a BGP withdrawal and DNS centralization locked out its own network.