tropicalthink — home
engineering

Building Practical High Availability Without Kubernetes

Infrastructure has a habit of becoming more complicated than the product it supports.

When we improved TropicalThink’s hosting setup, we wanted high availability, off-site recovery, automated failover, and useful operational visibility. The usual answers quickly led toward Kubernetes, managed load balancers, service meshes, and several new monthly subscriptions.

Those tools solve real problems, but they would have introduced more machinery than our products currently require. We are a small team operating a small estate. We wanted infrastructure we could understand, repair under pressure, and afford to keep running before traffic justified a larger platform.

This is not a universal alternative to Kubernetes or a claim that inexpensive servers become a cloud platform when connected together. It is a set of explicit trade-offs for our present scale.

What we run

Our application stack spans two inexpensive VPSs and an existing off-site server:

Topology

High-availability topology Two application nodes behind Cloudflare, running a PostgreSQL primary and synchronous standby, an off-site failover monitor, and encrypted backups in Cloudflare R2. Cloudflare R2 encrypted backups · off-site Cloudflare edge public ingress only PUBLIC TRAFFIC Node A Go API · cloudflared connector PostgreSQL primary Node B Go API · cloudflared connector PostgreSQL sync standby SYNC REPLICATION Failover monitor pg_auto_failover · off-site server HEALTH BACKUPS EVERY 6H

Cloudflare R2 holds encrypted backups outside the VPS estate. Staffroom, our internal operations application, receives health reports and shows the formation as one system.

The application nodes are deliberately symmetrical. If the primary fails, the monitor can promote the standby. When the failed server returns, it can rejoin as the new standby.

Following the writable database

The application needs neither a virtual IP nor a database load balancer. Its PostgreSQL connection configuration lists both database hosts and asks the driver for a writable server:

postgres://app@node-a:5433,node-b:5433/app?target_session_attrs=read-write

Connections drop during a switchover, then the pool reconnects to whichever node is writable. Application retries cover the short transition.

Failover

Database failover sequence The application connects to node A. Node A fails, the monitor promotes node B to primary, and the application reconnects to node B while node A rejoins as standby. Application pool · writable target Node A PostgreSQL primary Node B PostgreSQL sync standby Failover monitor watches both nodes

Cloudflare Tunnel replaces the public load balancer

Both application servers connect outward to the same Cloudflare Tunnel. If either connector disappears, Cloudflare can continue routing through the other. There is no floating public IP to move and no DNS record to change during an application-node failure.

Cloudflare Tunnel handles public ingress; it does not replace the network used for PostgreSQL replication and monitor communication. Those paths are separately restricted by host firewalls.

This makes Cloudflare a significant dependency. Two connectors protect us from losing a connector or VPS; they do not protect us from a Cloudflare outage, account problem, or configuration failure.

Tunnel ingress

Cloudflare Tunnel ingress Public traffic reaches the Cloudflare edge, which routes through either of two outbound tunnel connectors to node A or node B. If one connector fails, traffic continues through the other. Visitors public internet Cloudflare edge TLS · WAF · routing Node A tunnel connector Node B tunnel connector REROUTED THROUGH A

Why synchronous replication

A committed transaction must reach the synchronous standby before PostgreSQL confirms it to the application. This adds some write latency, but sharply reduces the chance of losing an acknowledged transaction if the primary disappears.

Availability and durability can pull in opposite directions. Without an eligible synchronous standby, PostgreSQL may pause writes rather than confirm data stored on only one machine. For our workload, preserving confirmed writes is the better default.

Replication is not a backup

A standby protects against machine failure. It does not protect against an accidental deletion that replicates immediately, compromised credentials, or corruption affecting the live estate.

We send encrypted backups to R2 every six hours and retain multiple recovery points. Backup operations are serialized so two machines cannot modify the same repository concurrently.

High availability keeps the service running after a node failure. Disaster recovery rebuilds it when the live copies cannot be trusted.

We tested object storage from both nodes, including retrieving the exact uploaded attachment through the surviving node’s credentials. A backup that has never restored data is only a hopeful theory.

The six-hour schedule is also a real limit. A disaster requiring an off-site restore may lose data created after the latest completed recovery point. Synchronous replication improves failover durability, but does not improve that backup recovery point. Continuous WAL archiving and a second independent backup destination remain future work.

Backup cadence · 6h RPO

Why the monitor lives elsewhere

pg_auto_failover uses a monitor to observe database nodes and coordinate promotion. We placed it outside the two-VPS application estate so a single provider or location failure is less likely to remove both the database and the machine coordinating recovery.

The primary keeps serving if the monitor is unavailable, although automatic failover cannot be coordinated safely. Instead of introducing another distributed system solely to cluster the monitor, we maintain and test an online replacement procedure.

The monitor is nevertheless a single coordination point. Its loss does not immediately stop the application, but it removes automatic promotion until we recover or replace it. Calling the entire system fully highly available would conceal that distinction.

Testing the failures we claim to support

We exercised the failure paths before relying on the design:

  • Controlled primary-to-standby switchovers
  • Loss of the primary, standby, and monitor
  • Split-brain prevention
  • Rejoining and completely rebuilding a standby
  • Backup repository contention
  • R2 upload and retrieval through either node
  • Public traffic through the remaining tunnel connector

These tests found real problems. Both backup timers could reach the same repository together. The repository correctly locked itself, but the second job initially treated that lock as a hard failure. We changed it to wait and retry for a bounded period.

These drills prove that specific scenarios worked under test conditions. They do not prove behavior under sustained production load, unusual network partitions, a provider-wide incident, or several simultaneous failures. We need to repeat them as versions, traffic, and topology change.

Staffroom sees systems, not isolated machines

A green CPU graph on one server says little about whether an application can survive a database failure. We changed Staffroom to aggregate host reports into service formations.

The Mitlist view now shows the current primary and standby, synchronous replication, byte lag, monitor reachability, healthy tunnel connectors, backup freshness, and every participating machine.

Staffroom can run HA checks, reconcile both nodes, back up the current primary, and request a controlled switchover. It is deliberately not a remote shell. Hosts pull commands over HTTPS, commands come from a closed set, and sensitive actions must pass checks in Staffroom and again on the host.

Where the design is weak

The architecture removes some single-machine failures, but it has clear limits:

  • Only two database nodes: losing one removes database redundancy. Another failure before recovery becomes a restore event.
  • A single monitor: the primary keeps serving without it, but automatic promotion cannot safely continue.
  • Shared external dependencies: both ingress paths depend on Cloudflare, while off-site recovery currently depends on R2.
  • Possible correlated failures: two VPSs offer limited protection if their provider, network, location, or administrative account fails as one unit.
  • No automatic capacity: a surviving node must already be able to carry the workload. Nothing creates capacity during an incident.
  • Synchronous writes can stop: our durability preference can reduce write availability when PostgreSQL cannot reach an eligible standby.
  • Manual enrollment: creating, hardening, and admitting a replacement VPS still requires operator work.
  • Small-team risk: knowledge concentration and the availability of an operator remain part of the recovery model.

Staffroom is also an operational dependency. If it or its D1 database is unavailable, the applications continue running, but the shared view and command queue disappear. Its information comes from periodic host reports and can be stale, which is why safety-critical commands must recheck conditions on the host.

Why we did not use Kubernetes

Kubernetes would provide scheduling, service discovery, reconciliation, networking, and a large controller ecosystem. It would also give us another distributed system to operate.

With only a few machines, its control plane, storage decisions, upgrades, and cluster-specific debugging would represent a significant share of our infrastructure. We would still need backups, database expertise, monitoring, and a recovery plan.

Our requirements are narrower: run known services, survive one application-node failure, promote PostgreSQL safely, keep origins private, store backups elsewhere, and rebuild machines from Git.

Systemd, Docker Compose, PostgreSQL, Cloudflare Tunnel, and a small reconciler cover those requirements well.

Kubernetes could become the more conservative choice if we need frequent dynamic scheduling, many independently deployed services, standardized autoscaling, stronger workload isolation, or a team already experienced in operating it. Avoiding it now does not make avoiding it forever a goal.

The cost of simplicity

We still own PostgreSQL upgrades, firewall rules, OS patching, failure tests, backup restoration, and the reconciler. There is no scheduler moving capacity around automatically.

Those constraints are acceptable at our current size because the system remains visible. We must revisit that judgement as the estate and team grow; familiarity is not a sufficient reason to keep an architecture after its assumptions stop being true.

What comes next

Our next step is repeatable enrollment: create a VPS, run one bootstrap command, approve its identity in Staffroom, assign it a role, and let Git-based reconciliation finish the work. Application capacity is the easier first target. Database nodes require stricter qualification because adding an unhealthy replica can reduce availability rather than improve it.

Likely improvements include a third database node in a separate failure domain, continuous WAL archiving, a second backup destination, automated restore verification, and measuring failover time under realistic load.

There is a risk that adding scheduling, networking, secret distribution, and lifecycle features would slowly produce a bespoke orchestrator. We intend to compare each proposed feature with established platforms instead of assuming custom code will remain simpler.

The goal is to automate operations we understand while keeping the properties we value: hosts pull, commands remain narrow, Git holds desired state, backups remain independent, and every redundancy claim is tested.

Our infrastructure is small because our needs are focused, not because reliability is unimportant. It gives us practical redundancy, automatic database failover, off-site recovery, and a foundation that can grow with our products.