Articles › High availability
Patroni and HAProxy: which health check to use (/primary, /replica, /leader and the rest)
HAProxy cannot see which PostgreSQL node is the primary. It asks Patroni, over HTTP, and Patroni answers with a status code: 200 means "send traffic here", 503 means "not me". Which endpoint you ask decides where your writes go during a failover, so it is worth getting right. Here is what each endpoint returns, which one to use for what, and the mistakes that cause writes to land on a replica.
The endpoints that matter
| Endpoint | Returns 200 when | Use it for |
|---|---|---|
/primary (also /read-write and /) | The node runs as primary and holds the leader lock | Read-write traffic |
/replica | The node is running, has the replica role, and does not carry the noloadbalance tag | Read-only traffic |
/replica?lag=1MB | The same, and replication lag is below the limit (bytes, or 16kB, 64MB, 1GB) | Read-only traffic that must not be stale |
/read-only | Like /replica, but the primary answers 200 too | Reads with the primary as a fallback |
/leader | The node holds the leader lock, whether it is a primary or the leader of a standby cluster | Rarely what you want for routing |
/synchronous, /asynchronous | The node is a synchronous or asynchronous standby | Reads from one kind of standby |
/health | PostgreSQL is up and running, whatever its role | Not for routing: every healthy node returns 200 |
/liveness, /readiness | The Patroni loop is healthy, or the node is ready and within acceptable lag | Kubernetes probes |
Everything else in HAProxy's configuration is plumbing around this table.
The configuration
This is the pattern our kit uses: two HAProxy listeners, one per kind of traffic, each checking a different endpoint on the Patroni REST API (port 8008 here) while sending clients to PostgreSQL itself (port 5432).
listen primary
bind *:5000
option httpchk
http-check send meth GET uri /primary
http-check expect status 200
default-server inter 2s fastinter 500ms downinter 1s fall 2 rise 1 on-marked-down shutdown-sessions
server pg1 10.0.0.11:5432 check port 8008
server pg2 10.0.0.12:5432 check port 8008
server pg3 10.0.0.13:5432 check port 8008
listen replicas
bind *:5001
balance roundrobin
option httpchk
http-check send meth GET uri /replica?lag=1MB
http-check expect status 200
default-server inter 2s fastinter 500ms downinter 1s fall 2 rise 1 on-marked-down shutdown-sessions
server pg1 10.0.0.11:5432 check port 8008
server pg2 10.0.0.12:5432 check port 8008
server pg3 10.0.0.13:5432 check port 8008
Applications write to port 5000 and read from port 5001. At any moment exactly one node answers 200 on /primary, so HAProxy sends writes to exactly one server.
Mistakes that send writes to the wrong node
Checking the PostgreSQL port instead of the REST API. A plain TCP check on 5432 is green on every node, replicas included. Clients then connect to a replica and get "cannot execute INSERT in a read-only transaction". Always check Patroni with check port 8008.
Using /health for the write pool. /health only says PostgreSQL is up. Every healthy node returns 200, so HAProxy balances your writes across all three.
Using /leader in a setup with a standby cluster. /leader is 200 for the leader lock holder, including the standby leader of a standby cluster. /primary is the one that means "accepts writes".
Forgetting on-marked-down shutdown-sessions. When a node is marked down, HAProxy only stops new connections by default. Sessions already open on a demoted primary stay open. This option closes them, so applications reconnect and land on the new primary instead of writing to the old one.
Slow checks. With inter 2s fall 2 rise 1, HAProxy notices a change in a few seconds. A check interval of 10 seconds would add its own delay on top of Patroni's. Detection time comes on top of the leader-key wait described in our failover measurements.
An unbounded read pool. /replica alone sends reads to a replica that has fallen far behind. Add ?lag= with a limit that suits your application. Exclude a replica from the read pool on purpose with the noloadbalance tag.
Check it from the shell
Before trusting HAProxy, ask each node yourself:
for h in 10.0.0.11 10.0.0.12 10.0.0.13; do
echo -n "$h primary: "; curl -s -o /dev/null -w '%{http_code}\n' http://$h:8008/primary
echo -n "$h replica: "; curl -s -o /dev/null -w '%{http_code}\n' "http://$h:8008/replica?lag=1MB"
done
On a healthy three-node cluster you expect one 200 on /primary and two on /replica. Anything else, such as two primaries or no primary, is what you want to find out before an incident, not during one.
During a failover there is a short window in which no node returns 200 on /primary. HAProxy marks all three down and the write port refuses connections until a new leader appears. Applications need a retry loop for that window.
Check the whole path, not only the endpoint
The REST API can be correct while HAProxy is misconfigured, and the other way round. The test that matters is: connect through the write port and ask the server whether it is in recovery.
psql -h 10.0.0.10 -p 5000 -Atc 'select pg_is_in_recovery()'
It must print f. If it prints t, writes are going to a replica. The free ha-check script runs this kind of routing check on the read-write and read-only ports in one pass, together with etcd, replication, archiving and backups: github.com/twinhullhq/ha-check.
The Twinhull kit ships these HAProxy and keepalived templates and a lab to try them on
The Twinhull HA Kit for PostgreSQL includes ready HAProxy and keepalived templates generated from one cluster file, plus a one-machine 3-node lab, so you can break the cluster and watch the health checks react before you do it in production.
See the kitTwinhull field notes
One tested Patroni fix or measurement a month, like the articles here. No spam; unsubscribe in one click.