Almost everything that holds your data today lives on more than one machine at once. And the moment several machines appear, an uncomfortable problem shows up too: they disagree. One server says the counter is 10, another says it is 12. This article explains the mechanism that keeps the truth from splitting in two: the quorum, that required majority that lets a distributed system keep working without going mad.
Two copies, two truths
The underlying problem is not speed or space: it is consistency. If you have two servers holding a copy of the same data and both accept writes on their own, as soon as two different operations arrive each one ends up with a different version. This is called state divergence, and it is the nightmare of any multi-node system.
The naive fix would be to require that all nodes agree before storing anything. That unanimous consensus works, but it has a brutal price: if a single node goes down or lags over the network, the whole system freezes. With just a little latency and a few machines, your operations start taking seconds. You need something in between: not all of them, but enough of them.
A quorum is just a number
The quorum defines how many nodes must acknowledge an operation for it to be valid. The most common form is the majority: with N nodes, the quorum is floor(N/2) + 1. With three servers, the quorum is 2; with five, it is 3.
The key is that two majorities always overlap. If two groups of voters each need more than half of the nodes, those two groups share at least one node. That overlap guarantees decisions chain together: each new decision “sees” at least one node that took part in the previous one, so nobody can write “over” what was already decided without anyone noticing.
How many failures you can take: the 2F+1 formula
How many machines do you need to tolerate F of them failing? The answer is 2F+1. To survive one failure (F=1) you need 3 nodes; two failures, 5 nodes. The reason: even if F nodes fail, the remainder is F+1, that is, the majority, which can keep deciding.
That surcharge is not waste: it is the mathematical cost of a majority surviving any partial failure. It is what lets a cluster lose one machine without stopping service, and why almost every production distributed system uses an odd number of control nodes.
Raft: electing a leader by majority
The Raft consensus algorithm uses the quorum for two things: electing a leader and replicating a log (the ordered sequence of changes that forms the system’s state). When the leader fails, the others fire their election timers and vote among themselves; the first to gather the majority becomes the new leader.
Every write decision travels as a log entry. The leader forwards it to everyone, and only commits it once a majority of nodes has copied it. That requirement is exactly the quorum: a committed entry is never lost, because it will always be held by a majority that still exists.
To avoid confusion between elections, votes carry a term (a number that grows with each election). A node only votes for a candidate from a more recent term, and if a leader discovers it has fallen behind, it steps down. That way the system avoids having two live leaders during the same era.
What happens when the network splits
The most delicate situation is a network partition: a cable is cut and the cluster is divided into two halves that cannot see each other. This is where the CAP theorem steps in: in that situation you must choose between consistency and availability. With a quorum, the choice is as clear as the majority: only the half holding the majority can keep writing; the other half drops to read-only or rejects requests.
That is exactly what prevents split brain, that disaster in which both halves write on their own and the system ends up with two irreconcilable truths. Without a quorum, that disaster is an accident waiting to happen; with one, it is a design decision that always falls on the same side.
The quorum also plays on reads
The quorum is not only for writing. Storage systems define two values: the write quorum W (how many nodes must confirm a write) and the read quorum R (how many must be read on a query). If W + R > N, any read touches at least one node holding the latest write, so you always read the most recent value.
That is the same majority-overlap principle applied to both operations. And it lets you fine-tune performance: with W=1, R=N you get lightning-fast writes and slow reads; with W=N, R=1, the opposite; with W=R=quorum, the classic balance.
The quorum in the real world
This idea is not textbook theory. etcd, the key-value store that powers Kubernetes, uses Raft and a majority quorum; Consul and ZooKeeper do the same at their core. Kafka replicates each partition to several in-sync replicas and only acknowledges a message once at least min.insync.replicas of them have copied it — a pragmatic quorum for a queue system.
Even PostgreSQL in synchronous mode lets you require several replicas to confirm before a transaction is committed. In every case the contract is the same: one node saying it is not enough; what is decided is what a majority supports. Accepting that a machine can fail is the first step; the majority is how that failure does not take your truth down with it.





