blog
Automated Database Failover: Why Homegrown High Availability Struggles At Scale
I’ve encountered DBAs and ops engineers who shared war stories of running production long before automation and modern tooling went mainstream. Experiencing sudden primary faults during sleeping hours, getting paged, logging in to squint over trace logs that show replication lag on all replicas and working out which one is furthest ahead, promoting it, editing an HAProxy config, reloading it, only to spend the next hour figuring out what to do with the old primary that just came back online and is accepting writes — tragic disaster!
If you are lucky, the procedure takes minutes and nobody outside the on-call rotation notices; if you are unlucky, you promoted a replica that was four seconds behind, and those four seconds contained payment records. The cost of downtime is easy to underestimate because it never shows up as one clear number, or simply gets buried across a dozen small things instead of one big number.
The fact is that there is lost revenue, the engineering time spent recovering instead of building, the SLA credit you owe your customers, and the erosion of trust that happens when your status page goes yellow for the third time in a quarter. For a team targeting 99.99% availability, the downtime allowance is approximately 52 minutes and 34 seconds over a 365-day year. One badly handled primary failure can consume a substantial share of that budget.
High availability is not optional at that point. What is worth questioning, is whether you should be building the failover machinery yourself. Most teams start by building it themselves, which is reasonable when you have one primary and one replica in one data center. The problem is that failover complexity does not grow linearly with your footprint. It grows with the number of topology and environment combinations you have to reason about, and that number multiplies fast.
Apart from that, it also forms complexity as the environment changes and the requirements and tooling goes bigger and bigger.
Failover Workflow Complexity
Before deciding whether a homegrown solution scales, you need a complete, thorough inventory of the steps. A reliable, working failover is not just a simple promotion of a replica, but a sequence of decisions, several of which present data loss risk if you get them wrong.
1. Node failure detection (heartbeats, thresholds)
This is the step people underestimate the most.
A single failed health check means nothing — a TCP connection can time out because the database is down, because the host is down, because a switch is being upgraded, because the kernel is thrashing under memory pressure, or because your monitoring host has a routing problem and the database is perfectly healthy. Every one of those looks identical from a single vantage point.
Getting heartbeat and threshold logic right means dealing with four things at once:
Multiple observers. If only your monitoring server can see the failure, the failure may be in your monitoring server. Real detection logic checks from the perspective of the replicas and the load balancers too. If the replicas can still stream from the primary and only your monitoring node cannot reach it, you almost certainly do not want to fail over.
Thresholds you can defend. Too aggressive and you fail over during a two second network blip, which is worse than the blip. Too conservative and your thirty second grace period becomes thirty seconds of downtime on every real failure. The number has to be justified against your network and your workload, not copied from a blog post.
Distinguishing down from unhealthy. A primary that is up but has a full disk, or is stuck behind a long running lock, or has hit max_connections, is not down. It is also not usable. Your logic has to decide whether these count as a failover trigger.
The load balancer’s view. The database is reachable from your monitoring system but not from your proxy layer. From the application’s perspective the cluster is unavailable; from your health check’s perspective everything is fine. Detection that only looks at the database misses this entirely.
2. Identifying the best candidate for promotion (which node has the latest data)
Once you accept that the primary is gone, you have to decide which node replaces it. Picking wrong here is how you lose transactions. This is also where the logic stops being portable between engines, because every engine tracks replication position differently. Let’s look at four engines that present four different answers to the same question. That is the first place a homegrown solution starts to strain, because the failover script you wrote for Postgres teaches you almost nothing about the one you need for MySQL, Galera Cluster, etc.
On PostgreSQL streaming replication, you compare the write ahead log positions on every surviving standby. That means sampling both the replay location and the receive location, because a standby may have received WAL it has not yet applied. The node with the highest LSN is your candidate. If you are running synchronous replication with a quorum, you also need to know which nodes were in the synchronous set, because a node outside that set may be arbitrarily behind.
On MySQL and MariaDB replication with GTID, you compare executed GTID sets. The more advanced replica is the obvious candidate, but you also have to check for errant transactions, meaning transactions that exist on a candidate but not on the old primary or the other replicas. Promoting a node with errant transactions will break replication for every other replica the moment they try to sync from it, and the failure will show up minutes later as a confusing duplicate key error rather than as an obvious failover problem.
If you are on file and position based replication instead of GTID, this comparison is considerably worse.
On Galera Cluster and MySQL Group Replication, the engine handles this for you. Galera uses quorum calculation and certification based replication, so node loss is resolved at the cluster level and there is no promotion to perform — a script that tries to promote a Galera node is not just unnecessary, it is dangerous. Group Replication elects a new primary in single primary mode on its own. Your job is not to pick a candidate but to
- understand what the cluster decided,
- reflect it in your topology map, and
- make sure the proxy layer knows which node is now the writer.
On MongoDB replica sets, the set holds its own election based on priority, optime, and votes. The engine does the promotion, but you still handle the consequences: connection strings and read preferences need to reflect the new primary, and in a sharded cluster the config servers and mongos routers need to converge on the new state.
Then come the data consistency checks, which is the part homegrown scripts skip most often because it demands the most engine specific knowledge. Having a candidate is not the same as having a cluster that can follow it. On PostgreSQL, you run pg_rewind against the standbys that diverged so their timelines reconcile with the newly promoted node. Skip it and you get standbys that will not attach, or worse, standbys that attach and silently diverge. On MySQL, you confirm the replicas you are about to repoint have no transactions the new primary lacks, create and grant the replication user, and repoint each replica in turn.
All of it serves one rule: exactly one writable node. Two nodes with read_only=OFF is a split brain in progress, and the damage is silent. Any serious failover implementation has to guarantee the old primary cannot accept a write, even if it comes back to life halfway through the promotion. That means fencing it, which in practice means being willing to shut down a database process that looks perfectly healthy.
3. Updating load balancers on DNS
Promotion means nothing until the application follows.
With a proxy such as ProxySQL, HAProxy, MaxScale, or PgBouncer, the configuration or backend health state has to be updated to reflect the new writer, and the change has to propagate to every proxy instance you run.
With DNS, you are at the mercy of TTLs and of resolvers and JVM clients that cache aggressively and ignore your TTL entirely.
With a virtual IP under Keepalived, you have to be sure the VIP cannot end up on two hosts at once.
Most real environments use some combination of all three, plus a set of application connection strings somebody hardcoded in 2021 and forgot about. Every one of those paths is a place where a technically successful promotion still reads as an outage to your users.
4. Reintroducing the failed node
The failover ends when the cluster is redundant again, not when the promotion succeeds. Until you rebuild the failed node, you are running without a spare, which means the next failure is a full outage rather than another failover — reintroduction is its own procedure.
The old node’s data is stale and, on PostgreSQL, likely on a divergent timeline, so it has to be rewound or rebuilt from a fresh base backup off the new primary. It has to come up read-only so it cannot accept writes during the process. Its replication then has to point at the new primary, which on a large data set can mean hours of streaming before you are back to full redundancy.
A single flowchart can help you visualize how the failover process works. The sample diagram below illustrates a simple workflow but each of these process can comprise complex components especially on a hybrid cloud or in a containerized environment.

DIY Approaches: Cron plus scripts for monitoring, manual intervention for DNS changes
Nobody sets out to build a fragile failover system. It represents an accumulation of decisions over time; version one is a cron job.
A shell script runs every minute, connects to the primary, and if the connection fails it writes to a log and sends an email. A human then does all the hard parts: reading the alert, comparing replication positions by hand, promoting, and editing the DNS record or the proxy config. This is honest, and it works, as long as someone is awake.
Version two automates the easy parts. The script now compares replication positions and runs the promotion itself. It lives in /usr/local/bin on the monitoring host with the replication credentials in a variable at the top, and it works on the topology it was written for. The DNS change is still manual, because nobody wants a cron job with permission to rewrite a production zone file.
Version three finally automates the traffic layer, so the script also SSHes into two HAProxy nodes and edits their configs with sed. This is where the special cases start. What happens if only one of the two accepts the change? What happens if the reload fails but the config is already written?
And version four adds a lock file, because someone discovered the cron job can fire while a previous invocation is still promoting.
Every version is a reasonable response to a real incident. That is exactly why the pattern is so common, and why it is rarely re-examined.
Each new node or environment multiplies the complexity
What makes this approach break down is not code quality; careful engineers write careful scripts. The problem is combinatorial, because every failover implementation is really a function of several independent variables:
- The database engine and version, e.g., Postgres 13 and Postgres 16 differ in how standby promotion can be triggered. PostgreSQL 16 removed
promote_trigger_file, so failover scripts that relied on trigger-file promotion need updating to usepg_ctl promoteorpg_promote(). MySQL 5.7 and 8.0 differ on Group Replication behaviour. - The topology. Primary with two replicas is not the same as primary with a cascading replica, which is not the same as a Galera cluster with an asynchronous replica hanging off it.
- The traffic layer. ProxySQL, HAProxy plus Keepalived, MaxScale, PgBouncer, a cloud load balancer, or plain DNS. Each has different update mechanics and different propagation behaviour.
- The environment. Bare metal, one cloud, another cloud, a Kubernetes cluster, and a staging environment configured slightly differently from production because of a decision no one remembers making.
A homegrown solution has to handle the product, not the sum, of these. Two engines, three topologies, and two traffic layers is twelve distinct code paths. Add a second region with different network characteristics and it is twenty four. Add one more replica per cluster and every position comparison, repointing loop, and proxy backend list must account for it. Each path needs to be written, and much more importantly, each path needs to be tested, because failover code is code you least want to discover is broken the moment you need it.
That testing burden is where most homegrown systems quietly fail. No one has time to induce twenty four different failure modes every quarter, so the paths get tested once at build time and then trusted forever. Meanwhile the environment drifts.
Someone upgrades a minor version, another adds a replica, and yet another changes a firewall rule. The script does not know any of this happened, and there is no signal that it has stopped working until the night it is needed.
There is an organizational cost too. The failover logic lives in one engineer’s head, in their bash, with their assumptions. They leave and you’re left with a system everyone is afraid to touch and nobody is willing to delete.
DIY failover degrades not because the engineering was bad, but because the surface area outgrows the team’s capacity to verify it.
N.B. A self-managed HA stack is not necessarily a homegrown failover script. Established tools such as Patroni already provide PostgreSQL failover coordination. The remaining question is who maintains, tests, and supports the complete service — including proxies, backups, monitoring, and recovery procedures. Building that stack can be a sound choice when your team has the expertise and capacity to own it. The case for ClusterControl is reducing the integration and lifecycle burden across supported databases and environments, not suggesting that reliable open-source HA is impossible.
Automated Failover with ClusterControl: The Work You No Longer Do
ClusterControl automates failover in one, streamlined workflow, supporting diverse types of database clusters, from MySQL and PostgreSQL, to MongoDB, ClickHouse and more. Users can operate mostly within its intuitive GUI but can alternatively operate through the s9s CLI for failing over manually. Let’s look at how auto-failover can be done with the click of a button.
In the list of clusters, simply enable auto-recovery either for cluster/node through this toggle switch:

or within the cluster dashboard itself as shown below:

Automatic failover becomes active as soon as cluster recovery is toggled on for that cluster; if toggled off node recovery will set off as well. However, toggling node recovery to off/on does not affect cluster recovery state.
What you get from there is one unified interface over several different underlying mechanisms. Where the database already elects a new primary on its own, as Galera Cluster, MySQL Group Replication, MongoDB replica sets, and Redis or Valkey Sentinel do through quorum and heartbeats, ClusterControl defers to the engine. It merely observes the outcome, updates the topology view, and adjusts monitoring and management for the new roles rather than competing with the election.
Where the database has no such mechanism, ClusterControl applies its own failover logic, following the vendor’s documented best practices. MySQL and MariaDB replication, along with PostgreSQL and TimescaleDB streaming replication, fall into that second category. Replication gives them redundancy, but neither ships with built-in fault tolerance or automatic recovery, so nothing in the stack will promote a replica unless something outside the database decides to.
Let’s briefly look at the difference between node and cluster recovery. Node recovery handles a single process dying.
If a database service, a ProxySQL instance, a HAProxy instance, a Keepalived daemon, a PgBouncer, or a Prometheus exporter stops when it should not have, ClusterControl waits about 30 seconds (default) for systemd to do its job. Then starts the service itself, and raises an alarm if it cannot. Notably, if you stopped the node deliberately through ClusterControl, it will not fight you and restart it.
Cluster recovery handles the topology being broken, which is the failover case. Both can be toggled independently per cluster.
A real failover event from ClusterControl’s perspective
The useful thing about an automated failover in ClusterControl is that it produces an auditable record of every decision. Logs can be found in /var/log/cmon.log or /var/log/cmon-audit.log. Here is what the sequence actually looks like per engine.
PostgreSQL and TimescaleDB streaming replication
- After 10 seconds of primary unreachability, it raises an alarm.
- After an additional 10 second waiting period, it starts the primary failover job.
- It samples replayLocation and receiveLocation on every surviving node to find the most advanced one.
- It promotes that node as the new primary.
- It stops the remaining replicas.
- It verifies the synchronization state using pg_rewind.
- It restarts the replicas pointing at the new primary.
If the promotion fails, the job aborts and raises an alarm rather than continuing into an unknown state. And when the old primary comes back online, ClusterControl forcibly shuts down its PostgreSQL service, because a former primary that restarts in writable mode can accept conflicting writes, and a writable stale primary is exactly the split brain you were trying to avoid.
That fencing behaviour is visible in the database log, the most direct evidence you can show someone who is skeptical that it happens:
2026-07-31 05:06:10.091 UTC [2392] LOG: database system is ready to accept connections
2026-07-31 05:06:27.696 UTC [2392] LOG: received fast shutdown request
2026-07-31 05:06:27.700 UTC [2392] LOG: aborting any active transactions
2026-07-31 05:06:27.703 UTC [2766] FATAL: terminating connection due to administrator command
2026-07-31 05:06:27.704 UTC [2758] FATAL: terminating connection due to administrator command
2026-07-31 05:06:27.709 UTC [2392] LOG: background worker "logical replication launcher" exited with exit code 1
2026-07-31 05:06:27.709 UTC [2414] LOG: shutting down
2026-07-31 05:06:27.735 UTC [2392] LOG: database system is shut down
The instance came up at 05:06:10 and was shut down 17 seconds later. That is the fence closing. Reincorporating that node as a replica is a deliberate user action, which is the right default: you should decide when a formerly diverged node rejoins your cluster.
MySQL and MariaDB replication
ClusterControl supports recovery for primary-replica and primary-primary setups with MySQL GTID or MariaDB GTID, as well as an asynchronous replica attached to a Galera cluster. The sequence:
- After 3 seconds of the primary being unreachable from multiple vantage points, including the ClusterControl host, the replicas, and the load balancers, it raises an alarm.
- It confirms at least one replica is reachable.
- It selects a promotion candidate.
- With GTID enabled, it calculates the probability of errant transactions.
- If none are detected, it promotes the candidate.
- It creates and grants the replication user.
- It repoints every replica that was following the old primary.
- It starts the replicas.
- It flushes logs on all nodes.
Two details are worth highlighting. First, all nodes are initially started with read_only=ON and super_read_only=ON regardless of role, and only one node is allowed read_only=OFF at a time. If two ever end up writable, ClusterControl sets both back to read only and requires a human to pick the real primary. That is the correct trade: a brief outage beats a silent split brain.
Second, if ProxySQL is in the cluster, ClusterControl also checks whether ProxySQL has any node available for writes. If the proxy layer has lost its writer even though ClusterControl can still reach every database node, the cluster is unavailable from the application’s point of view, and a failover is triggered. This is the failure mode homegrown health checks almost always miss.
Galera Cluster, Group Replication, and MongoDB
For engines with built in fault tolerance, ClusterControl deliberately does not try to run the election. Galera Cluster, MySQL NDB Cluster, MySQL Group Replication, MongoDB replica sets, Redis/Valkey Sentinel, Redis Cluster, and Elasticsearch all handle node loss natively through quorum, heartbeats, and role switching.
ClusterControl watches that happen, updates the topology view, and adjusts monitoring and management for the new roles; for example, recognising the new primary in a MongoDB replica set and making sure the proxy and exporter configuration follows it. This is the part that is hard to replicate with scripts. Knowing when to act is easy. Knowing when not to act, and instead to observe and reconcile, requires the tool to actually understand the engine.
Reliability: less data loss, faster recovery, less manual oversight
Three things change when the failover path is owned by a topology aware manager rather than by a script.
Data loss goes down because the promotion decision is based on measured replication position rather than on whichever node the on-call engineer checked first. On PostgreSQL, that is the highest position across replayLocation and receiveLocation, verified with pg_rewind before the replicas reattach. On MySQL, it is the executed GTID set plus an errant transaction check, which is the single most commonly skipped step in homegrown implementations and the one most likely to corrupt replication days after the incident is closed. Fencing the old primary closes the other data loss path, which is writes landing on a stale node nobody remembered to shut down.
Recovery gets faster; and, more usefully, it becomes predictable. The PostgreSQL path is roughly 20 seconds of alarm and grace period before the promotion job starts; the MySQL path raises its alarm after 3 seconds. Compare that to the human path, where the clock starts when a page is acknowledged rather than when the node fails, and where the difference between a 2 p.m. incident and a 2 a.m. incident is measured in tens of minutes. Predictability is what lets you put a number in an SLA and defend it.
Manual oversight drops to the decisions that genuinely need a human. Reattaching a diverged old primary, picking the real primary after a double-writable situation, and promoting a replica cluster in another data centre all still require someone to decide, and they should. Everything mechanical between detection and a working single-writer cluster does not.
One honest limitation is worth stating. Automatic recovery operates within a single cluster and does not extend to cluster-to-cluster replication. If your primary cluster in one data centre goes down entirely, promoting the replica cluster in the other data centre is a deliberate user action. For a full data centre failover you are making a business decision about accepting data loss, and that decision should not be made by a heartbeat timeout.
Laid side by side, the differences are less about speed than about which steps exist at all:
| Concern | Cron plus scripts | ClusterControl |
| Detection | Usually one vantage point | Controller, replicas, and load balancers |
| Candidate selection | Hand written per engine | Engine specific, LSN or GTID aware |
| Errant transaction check | Rarely implemented | Built into MySQL GTID recovery |
| Consistency verification | Often skipped | Pg_rewind on PostgreSQL, GTID checks on MySQL |
| Old primary fencing | Manual, easy to forget | Automatic, forced shutdown |
| Proxy and VIP update | Custom per proxy type | ProxySQL, HAProxy, MaxScale, Keepalived, PgBouncer |
| Native quorum engines | Risk of fighting the engine | Observes and reconciles instead |
| Cost of a new environment | Another code path to write and test | Configuration, not code |
| Audit trail | Whatever the script logged | Job log, alarms, notifications |
The row that matters most is the second to last one. With scripts, every new environment adds engineering work with no upper bound. With a topology aware manager, adding an environment is a deployment or an import, and the recovery logic you already trust applies to it.
Wrapping up
The bigger the environment, the more critical an automated failover system becomes. That is not a marketing line, it is arithmetic. One primary and one replica in one region is a problem a shell script can solve. Four engines across three regions with mixed proxy layers has more states than your team has time to test, and untested failover logic is indistinguishable from no failover logic until the night it matters.
The useful question is not whether you can build this. Of course you can. The question is whether you want your uptime to depend on code that only runs during your worst hours, that no one has exercised since it was written, and that no one remaining on the team fully understands.
Your next step: run a staging failover test
Do not take anyone’s word for this, including ours. Run the experiment.
Set up two staging clusters with the same topology, one behind your existing scripts and one managed by ClusterControl with cluster recovery enabled. Then intentionally kill the primary on both. Not a graceful shutdown, an actual kill -9 on the postmaster or mysqld; or, better yet, drop the network interface so you exercise your detection thresholds too.
Time it, and record more than the duration:
- How long until failover started? How long until the application could write again?
- Was any committed data lost? Verify it, do not assume it.
- Did every replica reattach, or did you have to rebuild one?
- Did the traffic layer follow, everywhere, including the proxy instance you forgot about?
- What happened when you brought the old primary back? Did it try to accept a write?
- How many manual steps were involved from failure to full redundancy?
Then repeat the test on a different engine, or with one extra replica in the topology, and see which of the two approaches needed changes to handle it. That second run is usually the one that settles the argument.
ClusterControl is free to try for 30 days, which is more than enough time to break a staging cluster on purpose several times over. If you want to see the failover path before you install anything, the recovery documentation walks through the exact sequence for each engine.
Database failover FAQ
What is automated database failover?
Automated database failover restores service after a failure by coordinating failure detection, selection or election of a replacement, and application access to the surviving database. In a single-primary topology, it must also prevent the old primary from continuing to accept conflicting writes. The exact mechanism depends on the database engine and topology.
Does automated failover guarantee zero data loss?
No. With asynchronous replication, transactions acknowledged by the primary may not have reached a surviving replica when it fails. A failover manager cannot recreate those missing transactions on the promoted replica. Data-loss exposure depends on replication mode, acknowledgment settings, replica availability, and the failure scenario—not automation alone.
Why must the old primary be fenced during failover?
In a single-primary system, an unreachable server is not necessarily a stopped server. A network partition can leave the old primary running while another node is promoted. Fencing prevents the old primary from accepting conflicting writes, protecting the cluster from two independent writable primaries.
How should you evaluate whether a homegrown failover still meets your needs?
Test more than replica promotion. Measure the time until applications can write again, whether acknowledged transactions survive, whether traffic reaches the correct node, and whether replicas return to a healthy topology. Revisit the tests after upgrades and topology changes. The decision is whether your team can maintain that verification discipline as the estate evolves.