Articles › High availability

Patroni and HAProxy: which health check to use (/primary, /replica, /leader and the rest)

Published 6 October 2026 · 5 min read · Checked against the Patroni 4.x REST API documentation and the HAProxy configuration in the Twinhull kit

HAProxy cannot see which PostgreSQL node is the primary. It asks Patroni, over HTTP, and Patroni answers with a status code: 200 means "send traffic here", 503 means "not me". Which endpoint you ask decides where your writes go during a failover, so it is worth getting right. Here is what each endpoint returns, which one to use for what, and the mistakes that cause writes to land on a replica.

The endpoints that matter

EndpointReturns 200 whenUse it for
/primary (also /read-write and /)The node runs as primary and holds the leader lockRead-write traffic
/replicaThe node is running, has the replica role, and does not carry the noloadbalance tagRead-only traffic
/replica?lag=1MBThe same, and replication lag is below the limit (bytes, or 16kB, 64MB, 1GB)Read-only traffic that must not be stale
/read-onlyLike /replica, but the primary answers 200 tooReads with the primary as a fallback
/leaderThe node holds the leader lock, whether it is a primary or the leader of a standby clusterRarely what you want for routing
/synchronous, /asynchronousThe node is a synchronous or asynchronous standbyReads from one kind of standby
/healthPostgreSQL is up and running, whatever its roleNot for routing: every healthy node returns 200
/liveness, /readinessThe Patroni loop is healthy, or the node is ready and within acceptable lagKubernetes probes

Everything else in HAProxy's configuration is plumbing around this table.

The configuration

This is the pattern our kit uses: two HAProxy listeners, one per kind of traffic, each checking a different endpoint on the Patroni REST API (port 8008 here) while sending clients to PostgreSQL itself (port 5432).

listen primary
    bind *:5000
    option httpchk
    http-check send meth GET uri /primary
    http-check expect status 200
    default-server inter 2s fastinter 500ms downinter 1s fall 2 rise 1 on-marked-down shutdown-sessions
    server pg1 10.0.0.11:5432 check port 8008
    server pg2 10.0.0.12:5432 check port 8008
    server pg3 10.0.0.13:5432 check port 8008

listen replicas
    bind *:5001
    balance roundrobin
    option httpchk
    http-check send meth GET uri /replica?lag=1MB
    http-check expect status 200
    default-server inter 2s fastinter 500ms downinter 1s fall 2 rise 1 on-marked-down shutdown-sessions
    server pg1 10.0.0.11:5432 check port 8008
    server pg2 10.0.0.12:5432 check port 8008
    server pg3 10.0.0.13:5432 check port 8008

Applications write to port 5000 and read from port 5001. At any moment exactly one node answers 200 on /primary, so HAProxy sends writes to exactly one server.

Mistakes that send writes to the wrong node

Checking the PostgreSQL port instead of the REST API. A plain TCP check on 5432 is green on every node, replicas included. Clients then connect to a replica and get "cannot execute INSERT in a read-only transaction". Always check Patroni with check port 8008.

Using /health for the write pool. /health only says PostgreSQL is up. Every healthy node returns 200, so HAProxy balances your writes across all three.

Using /leader in a setup with a standby cluster. /leader is 200 for the leader lock holder, including the standby leader of a standby cluster. /primary is the one that means "accepts writes".

Forgetting on-marked-down shutdown-sessions. When a node is marked down, HAProxy only stops new connections by default. Sessions already open on a demoted primary stay open. This option closes them, so applications reconnect and land on the new primary instead of writing to the old one.

Slow checks. With inter 2s fall 2 rise 1, HAProxy notices a change in a few seconds. A check interval of 10 seconds would add its own delay on top of Patroni's. Detection time comes on top of the leader-key wait described in our failover measurements.

An unbounded read pool. /replica alone sends reads to a replica that has fallen far behind. Add ?lag= with a limit that suits your application. Exclude a replica from the read pool on purpose with the noloadbalance tag.

Check it from the shell

Before trusting HAProxy, ask each node yourself:

for h in 10.0.0.11 10.0.0.12 10.0.0.13; do
  echo -n "$h primary: "; curl -s -o /dev/null -w '%{http_code}\n' http://$h:8008/primary
  echo -n "$h replica: "; curl -s -o /dev/null -w '%{http_code}\n' "http://$h:8008/replica?lag=1MB"
done

On a healthy three-node cluster you expect one 200 on /primary and two on /replica. Anything else, such as two primaries or no primary, is what you want to find out before an incident, not during one.

During a failover there is a short window in which no node returns 200 on /primary. HAProxy marks all three down and the write port refuses connections until a new leader appears. Applications need a retry loop for that window.

Check the whole path, not only the endpoint

The REST API can be correct while HAProxy is misconfigured, and the other way round. The test that matters is: connect through the write port and ask the server whether it is in recovery.

psql -h 10.0.0.10 -p 5000 -Atc 'select pg_is_in_recovery()'

It must print f. If it prints t, writes are going to a replica. The free ha-check script runs this kind of routing check on the read-write and read-only ports in one pass, together with etcd, replication, archiving and backups: github.com/twinhullhq/ha-check.

The Twinhull kit ships these HAProxy and keepalived templates and a lab to try them on

The Twinhull HA Kit for PostgreSQL includes ready HAProxy and keepalived templates generated from one cluster file, plus a one-machine 3-node lab, so you can break the cluster and watch the health checks react before you do it in production.

See the kit