RAID is not a backup. A mirror survives a dead disk and nothing else.
It does not survive a fat-fingered rm, a bad upgrade, or a fire. Real protection is layered.
The homelab follows the 3-2-1 rule: three copies of the data, on two kinds of media, with one
copy off-site. It builds this up as four layers, and each layer survives a bigger failure than the
last. ZFS makes this cheap: snapshots are near-free until data changes, and zfs send ships only
the blocks that moved.
Four layers of defense
Each layer is strictly stronger than the one before it. The first one barely counts as a backup at all. It is in the table only to make the point.What replicates, and what doesn’t
Not everything earns a second copy. Replication costs bandwidth, disk, and snapshot retention on the far side. It is reserved for data that is irreplaceable or is the system of record. Everything else gets local snapshots only. These are enough to undo a mistake, but they never get shipped across the wire.
The rule that decides which bucket a dataset lands in is simple:
How replication flows
Two always-on nodes replicate to each other on a nightly incremental schedule. Only changed blocks move, so even large datasets sync in seconds once seeded. A third node stays powered down most of the time. When it wakes, it pulls the latest snapshots from both always-on nodes, then shuts back off. That offline window is a feature, not a gap. A node that is powered off is air-gapped, so ransomware and a badzfs destroy can’t reach it. And because the cold node pulls rather than
being pushed to, a compromised primary has no standing credentials to corrupt the archive. The
powered-down copy is the “1” in 3-2-1.
The toolchain
Each concern maps to one well-worn open source tool. None of it is bespoke.
Snapshots and replication protect the filesystem; logical dumps protect the apps (a
database mid-transaction needs an app-consistent backup, not just a block snapshot). The two
are complementary, not redundant. A dedicated backup appliance was evaluated for this layer.
It was deliberately not adopted, because the dump-plus-replication pattern already covers the
gap without operating a second backup product.
A snapshot policy must not name a pool
A snapshot policy identifies its targets by dataset path, and a dataset path begins with the pool name. Writing that pool name into the policy by hand couples the policy to where the data lives today. When a guest is later moved to a different storage tier, a hand-written policy keeps snapshotting the old path. The old dataset usually still exists, since a tier move commonly retains the source. So snapshots keep being created and pruned on schedule. Snapshot counts and timestamps stay healthy while capturing nothing live. That failure is invisible to every count-based or age-based check, because the counts and ages are genuinely correct. They simply describe the wrong dataset. It is only detectable by comparing the policy’s target against the guest’s currently declared storage. The rule:- Resolve the dataset path from the guest’s declared storage at render time.
- Never write the pool name into the policy.
- If a guest’s storage cannot be resolved to a dataset (a directory-backed storage, for example, has no per-guest dataset), the render must fail loudly rather than emit a path or silently skip the guest.
What this connects to
Homelab
The hardware the pools run on.
ansible-proxmox
Where sanoid, syncoid, and the ZFS roles are defined.
tofu-proxmox
Declares the nodes, pools, and the guests that run on them.
Infrastructure overview
How the Proxmox stack fits the rest of the homelab.