Pages / #distributed-systems / #consensus
Raft consensus
Raft is a consensus algorithm that keeps one log identical on a cluster of servers: an elected leader takes every write and commits it once a majority has a copy.
Kill the leaderInteractive
- Published
- 2014, USENIX ATC (Best Paper)
- Votes to win with 5 servers
- 3
- Failures 5 servers survive
- 2
- etcd heartbeat
- 100 ms
- etcd election timeout
- 1000 ms
Three states, numbered terms
In words
Elections and terms
7/7
Raft cuts time into numbered terms, and each term starts with a leader election 1. A follower that hears no heartbeat for its election timeout increases the term counter, votes for itself and asks every other server for its vote 2. Each server votes once per term 3, a candidate with a majority becomes leader 4, and because the timeout is randomized, split votes are resolved quickly 5. etcd defaults to a 100 ms heartbeat and a 1000 ms election timeout 67.
7 of 7 quotes found in their sources
-
Raft (algorithm), Wikipedia en.wikipedia.org
Each term starts with a leader election.
Quote found in the source -
Raft (algorithm), Wikipedia en.wikipedia.org
It starts the election by increasing the term counter, voting for itself as new leader, and sending a message to all other servers requesting their vote.
Quote found in the source -
Raft (algorithm), Wikipedia en.wikipedia.org
A server will vote only once per term, on a first-come-first-served basis.
Quote found in the source -
Consensus, Consul documentation (HashiCorp) developer.hashicorp.com
If a candidate receives a quorum of votes, then it is promoted to leader.
Quote found in the source -
Raft (algorithm), Wikipedia en.wikipedia.org
Raft uses a randomized election timeout to ensure that split vote problems are resolved quickly.
Quote found in the source -
Tuning, etcd v3.5 documentation etcd.io
By default, etcd uses a 100ms heartbeat interval.
Quote found in the source -
Tuning, etcd v3.5 documentation etcd.io
By default, etcd uses a 1000ms election timeout.
Quote found in the source
Commit on a majority
5/5
Every write goes to the leader, which appends it to its log and replicates it to the followers 1. An entry is committed once it is durably stored on a majority, and only then is it applied 2. So 5 servers keep working with 2 down; with 3 down they stop making progress, but never return a wrong result 34. After a leader crash, the new leader repairs any mismatch by forcing the followers to duplicate its own log 5.
5 of 5 quotes found in their sources
-
Consensus, Consul documentation (HashiCorp) developer.hashicorp.com
A client can request that a leader append a new log entry.
Quote found in the source -
Consensus, Consul documentation (HashiCorp) developer.hashicorp.com
An entry is considered committed when it is durably stored on a quorum of nodes.
Quote found in the source -
The Raft Consensus Algorithm (raft.github.io) raft.github.io
a cluster of 5 servers can continue to operate even if 2 servers fail
Quote found in the source -
The Raft Consensus Algorithm (raft.github.io) raft.github.io
If more servers fail, they stop making progress (but will never return an incorrect result).
Quote found in the source -
Raft (algorithm), Wikipedia en.wikipedia.org
The new leader will then handle inconsistency by forcing the followers to duplicate its own log.
Quote found in the source
Built to be understood
5/5
Raft was published at USENIX ATC 2014 as an algorithm for managing a replicated log, equivalent to Paxos in fault tolerance and performance 12. It separates leader election, log replication and safety, and in a user study students learned it more easily than Paxos 34. The etcd Raft library powers etcd, Kubernetes and Docker Swarm, among others 5.
5 of 5 quotes found in their sources
-
In Search of an Understandable Consensus Algorithm, Ongaro and Ousterhout, USENIX ATC 2014 usenix.org
Raft is a consensus algorithm for managing a replicated log.
Quote found in the source -
The Raft Consensus Algorithm (raft.github.io) raft.github.io
It's equivalent to Paxos in fault-tolerance and performance.
Quote found in the source -
In Search of an Understandable Consensus Algorithm, Ongaro and Ousterhout, USENIX ATC 2014 usenix.org
Raft separates the key elements of consensus, such as leader election, log replication, and safety
Quote found in the source -
In Search of an Understandable Consensus Algorithm, Ongaro and Ousterhout, USENIX ATC 2014 usenix.org
Results from a user study demonstrate that Raft is easier for students to learn than Paxos.
Quote found in the source -
etcd-io/raft README (GitHub) github.com
It powers distributed systems such as etcd, Kubernetes, Docker Swarm
Quote found in the source