Avoiding MariaDB Galera Split-Brain in Two-Node ClustersAvoiding MariaDB Galera Split-Brain in Two-Node Clusters
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
The MariaDB Galera split-brain trap: You deployed a "high-availability" 2-node cluster, but a minor network glitch took down your entire production application.
MariaDB Galera Cluster is renowned for synchronous, multi-master replication with zero data loss (RPO=0). Seeing the benefits, engineering teams frequently spin up two database servers: "Node A and Node B replicate synchronously; if Node A dies, Node B takes over."
In distributed consensus engineering, however, a 2-node cluster is an architectural illusion that guarantees total downtime during network partitions.
Here is the exact consensus failure mode:
1. Quorum & Split-Brain Prevention: Galera uses strict majority voting (quorum) to guarantee consistency. In a cluster of $N$ nodes, a partition must control more than 50% of the votes ($\lfloor N/2 \rfloor + 1$) to remain in the "Primary" state and accept writes. 2. The Partition: In a 2-node cluster ($N=2$), the majority threshold requires at least 2 votes. If a transient network glitch interrupts communication between Node A and Node B, each node can only see itself (1 vote out of 2, which is 50%, not a majority). 3. The Total Lockout: Because neither node can prove it has a majority, both nodes immediately switch to ``non-primary`` state to prevent data divergence (split-brain). Both databases refuse all read and write queries, resulting in 100% application outage despite both servers being completely healthy.
Building true high-availability requires respecting distributed consensus invariants: • Odd-Numbered Node Quorum: Always deploy a minimum of 3 voting members across independent failure domains (availability zones or server racks). In a 3-node cluster, losing 1 node leaves 2 votes (66.7%), allowing the cluster to continue operating seamlessly without interruption. • Lightweight Galera Arbitrator (garbd): When hosting 3 full database servers is cost-prohibitive, deploy ``garbd`` on a lightweight third node or application server. The arbitrator does not store data or execute queries; it solely participates in consensus voting as an objective tie-breaker. • Automated State Transfer Hardening: Configure State Snapshot Transfers (SST) using ``mariabackup`` over encrypted TLS tunnels, ensuring new or recovering nodes re-join the cluster with zero table locking on donor nodes.
Never deploy a distributed database that locks itself up when a single cable hiccups.
Harden and scale your database infrastructure with our 2-week High-Availability Database Cluster Sprint on Contra: https://contra.com/s/QbN6svWo-high-availability-database-cluster-deployment-and-hardening
#Database #MariaDB #HighAvailability #Infrastructure #Architecture #Linux #DevOps #SRE
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started