Real output, not a mock-up
Failover drill report: hard crash
This is what failover-drill --mode crash writes after killing the leader's Patroni and PostgreSQL with SIGKILL while an application is writing. Nothing below is edited: it's the report from our 3-node lab (three PostgreSQL nodes on one VM, so the hosts are all 127.0.0.1).
Every committed row is checked afterwards. On your cluster the numbers will differ; measuring them is the point.
Failover drill report: PASS
| Cluster | lab |
| Run | drill-20260926-165412 (2026-09-26T16:54:51+03:00) |
| Mode | crash |
| Replication | asynchronous |
| Old leader → new leader | pg1 → pg3 |
| Timeline | 1 → 2 |
Results
| Metric | Value |
|---|---|
| New leader routable via HAProxy after | 24.4 s |
| Write downtime (first failed → first successful write) | 24.2 s |
| Longest gap between successful commits | 24.4 s |
| Writes attempted / committed / failed | 129 / 53 / 76 |
| Committed transactions lost | 0 |
| Old leader back as streaming replica after | 33.8 s |
Notes
- No loss this time, but async replication does not guarantee it.
Cluster after the drill
+ Cluster: lab (7689841837669857905) -----------+----+-------------+-----+------------+-----+
| Member | Host | Role | State | TL | Receive LSN | Lag | Replay LSN | Lag |
+--------+----------------+---------+-----------+----+-------------+-----+------------+-----+
| pg1 | 127.0.0.1:5432 | Replica | streaming | 2 | 0/603C8B8 | 0 | 0/603C8B8 | 0 |
| pg2 | 127.0.0.1:5433 | Replica | streaming | 2 | 0/603C598 | 0 | 0/603C598 | 0 |
| pg3 | 127.0.0.1:5434 | Leader | running | 2 | | | | |
+--------+----------------+---------+-----------+----+-------------+-----+------------+-----+
How to read it
- Write downtime is what your users feel: from the first failed write to the first successful write on the new leader.
- Writes failed are attempts made while there was no leader; a retrying connection pool absorbs them. Committed transactions lost is the number that matters, and it's checked row by row.
- About 24 seconds is expected with Patroni's default
ttlof 30 seconds: a crashed leader can't release its lock, so it has to expire. Why, and what tuning buys you. - Old leader back as streaming replica shows that
pg_rewindworked and the node rejoined by itself.
Run the same drill on your cluster
The kit includes failover-drill (switchover, clean stop and crash modes), ha-check, the configuration templates, a 3-node lab and the 10-chapter guide.
Twinhull field notes
One tested Patroni fix or measurement a month, like the articles here. No spam; unsubscribe in one click.