Skip to main content
Status: live. Failover management, streaming replication, and both backup tiers run today; the on-site restore drill passes.
Everything else is rebuildable from source. The data is not. So the data gets the protection, and the rest gets to be disposable.
The homelab’s design philosophy is deliberate disposability: hosts and services are reconstructed from code, not nursed. That only works if the irreplaceable data sits behind real protection while everything around it stays cheap to rebuild. The shared app database is that irreplaceable layer, so it carries two independent resilience tiers.

High availability

Failover management is live. The cluster manages the database containers as highly available resources. So losing a node triggers a managed restart on a surviving node instead of a manual recovery. An anti-affinity rule keeps the two database containers apart, so no single node can take both. A primary and a standby database stay in sync through continuous streaming replication. The former primary now runs as a warm standby: it applies every write the primary ships and reports a healthy streaming state from both ends. The apps were cut over to the new primary with a single switch. This was verified by matching table counts and live client connections before the old primary was demoted.

Layered backups

Backups are layered so no single failure, whether hardware, mistake, or malice, can take the data with it. The on-site tier is live and rehearsed. The database ships every write to it continuously. It also takes a full base backup weekly with incrementals in between. So a restore can land on any point in time, rather than on the last nightly snapshot. A restore drill passes: the databases come back queryable. The off-site copy is meant to be identical to the on-site one: same layout, same names. So a restore reads the same paths either way and only swaps which store it points at. The off-site copy is rebuilt and verified at the object level: every object matches the on-site tier. The check asks for the specific manifest file a restore needs, rather than trusting a directory listing. Restore rehearsal covers the on-site tier. The database can also write to both stores natively as twin repositories, which retires the mirror step entirely. That mode is built and tested but not yet switched on. A directory listing proves the listing matched. It does not prove a copy is restorable. A check built on the same index the copy itself used can report perfect parity while still missing the one file a restore actually needs. The reliable check asks for that specific artifact directly: the manifest a restore reads first, not an inventory comparison of what’s present. This tier sits alongside the host-level ZFS backup and replication that protects the guests themselves. Together they cover both the box and the data inside it.