The Ceph rebalance death spiral: When a single failed disk takes down an entire enterprise virtua...The Ceph rebalance death spiral: When a single failed disk takes down an entire enterprise virtua...
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
The Ceph rebalance death spiral: When a single failed disk takes down an entire enterprise virtualization cluster.
Ceph is one of the most resilient software-defined storage architectures ever engineered. When designed properly, it handles drive dropouts, host crashes, and rack outages without missing a single block I/O.
When deployed without strict physical network isolation, however, Ceph harbors a catastrophic failure mode: the cascading rebalance death spiral.
Here is the exact anatomy of the failure we forensically triaged on a 150-host private cloud cluster:
1. The Single Pipe Bottleneck: The cluster was configured with client VM traffic (public network) and internal Ceph replication/backfill traffic sharing the same physical 10GbE network interfaces. 2. The Trigger: A routine NVMe drive failed, transitioning one OSD to ``down/out``. Ceph immediately did what it was designed to do: calculate new placement group mappings and replicate degraded objects across the surviving nodes. 3. The Congestion Collapse: The flood of backfill data saturated the shared NICs. Public client traffic was choked out. OSD heartbeat pings across the cluster exceeded ``osd_heartbeat_grace`` (20 seconds). 4. The Death Spiral: The Ceph monitor quorum marked 8 more OSDs as ``down`` due to missed heartbeats. That triggered even more backfill replication across an already asphyxiated network, freezing all client KVM disk I/O and locking up hundreds of production VMs.
The fix required three non-negotiable architectural remediations: • Physical segregation of public client traffic and dedicated cluster backfill traffic across bonded 25GbE interfaces. • Strict MTU 9000 (Jumbo Frames) end-to-end across all storage switches and host interfaces to eliminate CPU packet fragmentation overhead. • Precise QoS limits on backfill and recovery I/O (``osd_max_backfills = 1``, ``osd_recovery_max_active = 2``) ensuring client I/O always retains 80%+ priority during cluster healing.
If your private cloud storage has never been tested under full-load drive loss scenarios, you don't have high availability—you have a ticking time bomb.
Migrate and harden your virtualization storage with our VMware Virtualization Rescue & Ceph Migration Sprint on Contra: https://contra.com/s/FocA2qgU-v-mware-virtualization-rescue-sovereign-kvm-and-ceph-migration
#Infrastructure #Linux #DevOps #SRE #Architecture #Performance #Cloud
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started