zfs-impermanence-rollback¶
Modules
Wipe-on-boot for ZFS: declare which datasets get rolled back to a blank
snapshot inside the initrd, and get the systemd ordering, the neededForBoot
guard rail and the loud-failure behaviour that make the wipe an actual
guarantee instead of a hope.
The interesting part of this recipe is not rolling back /. Everyone gets
that right, because there is exactly one obvious unit to order against. The
interesting part is the second dataset — usually /home — where the
obvious answer is wrong in a way that produces a machine which boots fine,
looks right, passes every check you think to run, and quietly keeps state
forever.
zfsWipeOnBoot = {
enable = true;
datasets = {
root = { dataset = "rpool/local/root"; mountPoint = "/"; };
home = { dataset = "rpool/local/home"; mountPoint = "/home"; };
};
};
The problem¶
"Erase your darlings" impermanence has two halves:
- Put the state you want back somewhere durable. That is
nix-community/impermanence'senvironment.persistence.<root>— bind mounts from a persistent dataset onto an ephemeral root. - Destroy everything else on every boot. That is a
zfs rollbackin the initrd, before anything mounts the dataset.
nixpkgs implements neither, and the impermanence flake implements only the
first. There is no fileSystems.<name>.wipeOnBoot, no boot.zfs.rollback*,
nothing. nixos/modules/tasks/filesystems/zfs.nix (~1500 lines, covering
import, key loading, auto-snapshot, scrub, trim, ZED, expand) has no notion of
rolling anything back. So every impermanence host in the world carries a
hand-written copy of the same twelve-line unit, and the copies differ in ways
their authors did not intend.
Half of them are also still written against scripted stage 1:
That form has no ordering semantics whatsoever — it is a shell fragment
concatenated into one big script — and under the systemd initrd it is a
removed option: nixos/modules/system/boot/systemd/initrd.nix lists
postDeviceCommands in its obsoleteOpt block with the message "systemd stage
1 does not support boot.initrd.postDeviceCommands". A host that flips
boot.initrd.systemd.enable = true and forgets to port its rollback does not
get a warning about a disabled wipe; it gets an evaluation error, which is
the good case. A host that ports it incorrectly gets no error at all.
This module handles the second half only. Pair it with the impermanence flake, or with plain bind mounts, for the first.
Trap 1 — the anchor must be sysroot.mount, even for /home¶
This is the whole reason the recipe exists.
Write the root wipe and there is one candidate unit and it is right:
boot.initrd.systemd.services.rollback-root = {
wantedBy = [ "initrd.target" ];
after = [ "zfs-import-rpool.service" ];
before = [ "sysroot.mount" ]; # <-- the anchor
unitConfig.DefaultDependencies = "no";
serviceConfig = { Type = "oneshot"; RemainAfterExit = true; };
script = "zfs rollback -r rpool/local/root@blank";
};
Now add /home. The symmetric-looking edit is to copy the unit and change
sysroot.mount to the mount unit for /home:
It even works, on the host you tested it on, today. Here is why it is not a guarantee.
What the initrd actually does¶
neededForBoot filesystems are written into a separate fstab that is handed to
systemd in the initrd (nixos/modules/tasks/filesystems.nix):
initrdFstab = pkgs.writeText "initrd-fstab" (
makeFstabEntries (filter utils.fsNeededForBoot fileSystems) { }
);
...
boot.initrd.systemd.storePaths = [ initrdFstab ];
boot.initrd.systemd.managerEnvironment.SYSTEMD_SYSROOT_FSTAB = initrdFstab;
boot.initrd.systemd.services.initrd-parse-etc.environment.SYSTEMD_SYSROOT_FSTAB = initrdFstab;
systemd-fstab-generator turns each entry into a /sysroot-prefixed mount
unit: / → sysroot.mount, /home → sysroot-home.mount, /var/log →
sysroot-var-log.mount. The two are not peers:
sysroot.mountis pulled byinitrd-root-fs.target.- everything else is pulled by
initrd-fs.target, which isAfter=initrd-parse-etc.service, which isRequires=/After=initrd-root-fs.target. - and, independently, systemd.mount(5) states: "If a mount unit is beneath
another mount unit in the file system hierarchy, both a requirement
dependency and an ordering dependency between both units are created
automatically."
sysroot-home.mountis beneathsysroot.mount, so it gainsRequires=sysroot.mountandAfter=sysroot.mountfor free.
So sysroot.mount is a strictly earlier, strictly more certain point in
the initrd timeline than any other sysroot-*.mount. Ordering the wipe before
sysroot.mount transitively orders it before every one of them. That is the
"ordering between the two mount units" that the anchor leans on, and it is why
before = [ "sysroot.mount" ] is the correct answer for a dataset mounted at
/home just as much as for one mounted at /.
Why the narrow anchor fails silently¶
Before= is not a requirement. systemd.unit(5): ordering is orthogonal to
Requires=/Wants=, and a unit that is not part of the transaction imposes no
ordering at all. A Before= naming a unit that does not exist is not an
error, not a warning, not a failed assertion. It is a no-op, and systemd logs
nothing.
sysroot-home.mount does not exist unless /home is neededForBoot. And
neededForBoot is a separate line, in a separate file, usually written by
whoever set up disko — not by whoever wrote the rollback unit. Drop it, move
/home to a different mount point, mkForce the fileSystems entry from a
hardware module, split /home into per-user datasets — any of those removes
the unit, and the Before= evaporates with it.
At that point the wipe service has DefaultDependencies=no (mandatory: with
default dependencies a service is After=basic.target, far too late) and one
surviving constraint, After=zfs-import-rpool.service. Nothing orders it
against the rest of the boot. Two things then happen, both quiet:
- It can run after stage 2 has already mounted
/home— and a mounted dataset cannot be rolled back while anything holds it open:cannot rollback 'rpool/local/home': mountpoint or dataset is busy. - More likely, it never runs at all.
initrd-cleanup.serviceis literallyExecStart=systemctl --no-block isolate initrd-switch-root.target, andisolatestops every unit that is not a dependency of the isolated target. AWantedBy=initrd.target,DefaultDependencies=noservice with no relationship toinitrd-switch-root.targetis exactly that: it gets a stop job and is torn down, started or not.
Either way initrd.target only Wants= its filesystems, the boot proceeds,
you land at a login prompt, /home is intact, and that is what you asked
for as far as anything can tell. No unit is red. No message mentions the
wipe. The threat model is gone and the only symptom is that last month's
browser profile is still there — which reads as a feature.
sysroot.mount cannot disappear. Every bootable Linux system has a root
filesystem, and / is unconditionally in utils.pathsNeededForBoot
(nixos/lib/utils.nix), so its mount unit is in the initrd regardless of what
anyone writes in fileSystems.
This module emits both, sysroot.mount first:
[Unit]
After=zfs-import-rpool.service
Before=sysroot.mount sysroot-home.mount
DefaultDependencies=no
The dataset's own mount unit is redundant given the anchor. It is there so the next reader can see which mount this unit is protecting without deriving the escaped name in their head.
Trap 2 — neededForBoot is not a performance hint¶
Because Trap 1 hinges on it, this module sets it rather than hoping:
and then asserts on the effective value anyway, so a mkForce from a
hardware module elsewhere is caught at eval:
zfsWipeOnBoot.datasets.home wipes rpool/local/home mounted at /home, but that
filesystem is not neededForBoot.
Its `sysroot-home.mount` unit therefore does not exist in the initrd, the mount
happens in stage 2 instead, and the wipe is unordered with respect to it.
The check is utils.fsNeededForBoot, i.e. fs.neededForBoot || elem
fs.mountPoint pathsNeededForBoot, not the raw option — so / never trips it
spuriously.
One shape defeats the setting while leaving the assertion intact, which is
the correct trade: if a host writes the whole filesystem entry with mkForce,
fileSystems."/home" = lib.mkForce {
device = "rpool/local/home";
fsType = "zfs";
options = [ "zfsutil" ];
neededForBoot = true; # <-- you now own this line
};
then priority 50 replaces the submodule wholesale and this module's
neededForBoot = true (priority 100) is discarded. That is fine as long as the
forced value says true; if it does not, the assertion fires at eval instead
of the wipe failing at 3 a.m.
Setting neededForBoot has a visible side effect worth knowing: nixpkgs adds
x-initrd.mount to the filesystem's options
(nixos/modules/tasks/filesystems.nix, the config.options mkMerge). If you
diff a host's fileSystems."/home".options before and after adopting this
module, that is the expected change.
Trap 3 — a failed wipe is a successful boot¶
initrd.target declares Wants=initrd-root-fs.target initrd-root-device.target
initrd-fs.target initrd-usr-fs.target initrd-parse-etc.service — Wants=,
not Requires=. A wantedBy = [ "initrd.target" ] oneshot that exits non-zero
does not stop anything. zfs rollback failing because the snapshot was never
created is therefore indistinguishable, from the outside, from it succeeding.
So this module defaults failHard = true, which emits the same pair systemd's
own initrd units use:
replace-irreversibly matters: without it the pending boot transaction can
still complete around the emergency job. This is exactly what
initrd-parse-etc.service, initrd-fs.target and initrd-cleanup.service
ship with upstream.
And because "the snapshot does not exist" is by far the most common cause, the
generated script checks first and says so, instead of leaving you to decode
cannot open 'rpool/local/home@blank': dataset does not exist from a
half-second of scrollback:
zfs-wipe-on-boot: rpool/local/home@blank does not exist. Refusing to boot with
state that was supposed to be discarded. Create it with:
zfs snapshot rpool/local/home@blank
failHard = false on hosts you cannot reach. emergency.target in the
initrd is a serial console prompt. If the host has neither
boot.initrd.systemd.emergencyAccess = true nor initrd SSH (see
remote-luks-unlock), a hard fail is a brick that
needs someone with physical access. On such hosts prefer failHard = false and
a boot-time alert, and accept that you must monitor for the failure yourself.
Trap 4 — -r silently eats every snapshot of that dataset, every boot¶
zfs rollback refuses by default to roll back to anything but the most recent
snapshot:
which means an unattended wipe must pass -r — the moment any snapshot
timer, replication job or zfs-auto-snapshot touches the dataset, a rollback
without -r fails on every subsequent boot. So recursive defaults to true.
The consequence, from zfs-rollback(8): -r "Destroy any snapshots and
bookmarks more recent than the one specified." On a dataset carrying
com.sun:auto-snapshot=true that is: destroy the entire snapshot history of
that dataset on every single boot, without a word.
Disko templates make this easy to get wrong, because the sensible-looking
default for a home dataset is to snapshot it:
"safe/home" = {
type = "zfs_fs";
mountpoint = "/home";
options."com.sun:auto-snapshot" = "true"; # correct for a PERSISTENT /home
};
If /home is wiped on boot, that property must be "false", and the durable
copies live on the persistent dataset instead. Keep the two shapes clearly
separate in your disko templates and never let a "wipe /home" host inherit the
"keep /home" dataset options.
ZFS dataset properties are live state, not declarative config: disko's
options apply at pool-creation time and never re-run. A module cannot detect
this at eval, so it is documentation and a habit, not an assertion. Check it
by hand:
Trap 5 — the blank snapshot must predate the first boot, idempotently¶
@blank has to exist before the wipe first runs, which means at install time,
which means a disko postCreateHook:
"local/home" = {
type = "zfs_fs";
options = { mountpoint = "/home"; canmount = "noauto"; };
postCreateHook = "zfs snapshot rpool/local/home@blank";
};
Write it idempotently. postCreateHook re-runs on any repeat of the create
step, and a second zfs snapshot of an existing name fails the whole disko
run:
postCreateHook = ''
zfs list -t snapshot -H -o name | grep -qE '^rpool/local/home@blank$' \
|| zfs snapshot rpool/local/home@blank
'';
For a dataset added to an existing host there is no install step to hook, so
this module offers onMissingSnapshot = "create": it snapshots the current
contents under that name and continues, logging clearly that this boot kept its
state. Use it for exactly one boot, then set it back to "fail" — left on, it
turns "the snapshot is missing" from a loud stop into an automatic
re-baselining of whatever happened to be on disk.
Trap 6 — make sure nothing else mounts the dataset first¶
The wipe races anything that mounts the dataset outside the mount units you ordered against. Two upstream behaviours are on your side, and both are worth knowing so you do not accidentally opt out:
- Import does not mount. The initrd import runs
zpool import -d … -N …(nixos/modules/tasks/filesystems/zfs.nix);-Nis "import without mounting". Nothing in the pool is mounted by the import itself. - The import unit is ordered before every one of the pool's mounts. The
same file computes
getPoolMounts— the/sysroot-prefixed, escaped mount unit name for each of that pool'sneededForBootfilesystems — and sets bothrequiredByandbeforeto it. That is why this module derives itsAfter=zfs-import-<pool>.servicefrom the dataset's own pool prefix rather than from an option you could get wrong.
What can still bite you is stage 2. zfs-mount.service mounts datasets with a
real mountpoint property and canmount=on. For a wiped dataset that is at
best a double mount and at worst a mount of the pre-rollback state over the
top. Two ways out, both used in the wild:
mountpoint = "legacy"— ZFS will not mount it; only the fstab/mount unit does. Simplest.mountpoint = "/home"+canmount = "noauto"+ mount optionzfsutil— keeps the property useful forzfs mountby hand whilezfs mount -askips it. Hosts that go further disable the unit outright withsystemd.services.zfs-mount.enable = false.
Pick one per host and be consistent; mixing them is how a dataset ends up
mounted twice with the wipe applied to the copy nobody is using. This module
warns when fileSystems.<mountPoint>.device is not the dataset it was told to
roll back, which catches the most common form of that mistake.
Trap 7 — encryption changes the ordering, not the anchor¶
With the pool inside LUKS, the import unit needs its own dependency on the cryptsetup unit; the wipe inherits the ordering transitively and needs no change:
boot.initrd.systemd.services."zfs-import-rpool" = {
after = [ "systemd-cryptsetup@crypted.service" ];
requires = [ "systemd-cryptsetup@crypted.service" ];
};
With ZFS native encryption the key load happens inside the import unit
itself, so there is nothing extra to order — see
zfs-native-encryption-keys for how the key
gets into the initrd in the first place. Either way the per-dataset after
option here is for units the import depends on, not for the wipe.
Trap 8 — what must be on the persistent dataset, or the wipe rotates it¶
A wipe that includes / destroys machine identity unless it is persisted
explicitly. The two that hurt:
/etc/machine-id— regenerated every boot. Breaks the persistent journal (/var/log/journal/<machine-id>/becomes a new directory each boot, and the old ones are never read), systemd'sConditionFirstBoot, and anything keyed on it.- SSH host keys — regenerated every boot, so every client gets a host-key
mismatch on every reboot, and the fleet-wide muscle-memory response to that
is to delete the
known_hostsline, which is exactly the reflex a real man-in-the-middle needs you to have. Point them at the persistent dataset withservices.openssh.hostKeysand check the paths actually resolve there.
Also persist /var/lib/nixos (UID/GID allocations — without it, dynamic users
renumber and file ownership on persisted data drifts) and /var/lib/systemd.
For a wiped /home specifically, remember that per-user persistence
bind-mounts into a home directory that the wipe just emptied. The user's home
must be recreated (it is, by systemd-tmpfiles / users.users.<n>.home) with
the right ownership before the bind mounts land, and anything not listed is
gone. The failure mode is not data loss you notice — it is a login that
silently resets a setting you changed three weeks ago.
Usage¶
{
imports = [ ./zfs-impermanence-rollback ];
boot.initrd.systemd.enable = true;
zfsWipeOnBoot = {
enable = true;
datasets = {
root = { dataset = "rpool/local/root"; mountPoint = "/"; };
home = { dataset = "rpool/local/home"; mountPoint = "/home"; };
};
};
fileSystems."/persist".neededForBoot = true;
environment.persistence."/persist" = {
hideMounts = true;
directories = [ "/var/lib/nixos" "/var/lib/systemd" ];
files = [ "/etc/machine-id" ];
};
services.openssh.hostKeys = lib.mkForce [
{ path = "/persist/etc/ssh/ssh_host_ed25519_key"; type = "ed25519"; }
];
}
Root inside LUKS, headless, no console access:
zfsWipeOnBoot = {
enable = true;
datasets.root = {
dataset = "rpool/local/root";
mountPoint = "/";
failHard = false; # see Trap 3
after = [ "systemd-cryptsetup@crypted.service" ];
};
};
Adding a wiped /home to a host that already has one, for a single boot:
zfsWipeOnBoot.datasets.home = {
dataset = "rpool/local/home";
mountPoint = "/home";
onMissingSnapshot = "create"; # then change back to "fail"
};
Options¶
| Option | Default | Effect |
|---|---|---|
zfsWipeOnBoot.enable |
false |
Nothing is emitted until this is on. |
zfsWipeOnBoot.package |
config.boot.zfs.package |
zfs binary put on the unit's PATH. |
zfsWipeOnBoot.snapshot |
"blank" |
Default snapshot name for every entry. |
zfsWipeOnBoot.namePrefix |
"zfs-wipe" |
Units are <prefix>-<key>.service. |
zfsWipeOnBoot.enforceNeededForBoot |
true |
Sets fileSystems.<mountPoint>.neededForBoot = true. See Trap 2. |
zfsWipeOnBoot.datasets |
{ } |
Attrset of entries, keyed by unit-name suffix. |
…datasets.<n>.enable |
true |
Turn one entry off without deleting it. |
…datasets.<n>.dataset |
required | Full dataset name; its pool prefix picks the import unit. |
…datasets.<n>.snapshot |
zfsWipeOnBoot.snapshot |
Snapshot name without the @. |
…datasets.<n>.mountPoint |
null |
Enables the neededForBoot guard rail and assertions. |
…datasets.<n>.recursive |
true |
zfs rollback -r. See Trap 4. |
…datasets.<n>.onMissingSnapshot |
"fail" |
fail / create / ignore. See Trap 5. |
…datasets.<n>.failHard |
true |
OnFailure=emergency.target. See Trap 3. |
…datasets.<n>.after |
[ ] |
Extra After=, e.g. a cryptsetup unit. |
…datasets.<n>.requires |
[ ] |
Extra Requires=. |
Emitted per entry:
[Unit]
After=zfs-import-<pool>.service
Before=sysroot.mount <escaped mount unit>
DefaultDependencies=no
OnFailure=emergency.target
OnFailureJobMode=replace-irreversibly
[Service]
Type=oneshot
RemainAfterExit=true
Verifying the wipe actually happens¶
Do not trust "it booted". Three checks, in increasing strength:
-
The canary.
touch ~/canary-do-not-persist, reboot, look. Thirty seconds, and it catches every failure mode in this document. -
The unit ran, and ran early. The initrd journal is flushed into the main journal at switch-root, so:
journalctl -b -o short-monotonic -u zfs-wipe-home.service
journalctl -b -o short-monotonic --grep 'sysroot|zfs-wipe'
The Finished …zfs-wipe-home.service line must precede
Mounted /sysroot/home. If neither line is present at all, the unit was
never started — that is Trap 1, and it is the answer more often than
anything else. Add rd.systemd.log_level=debug to boot.kernelParams for
one boot if the ordering is not obvious from the timestamps.
Caveat: on a host whose /var/log is itself wiped and not persisted, this
evidence is destroyed by the next reboot. Persist /var/log (or a journal
directory) before you start debugging boot ordering.
- The dataset is genuinely at the snapshot. Right after boot:
zfs get -H -o value written rpool/local/home # bytes changed since @blank
zfs list -t snapshot -o name,creation rpool/local/home
written should be small and growing from zero this boot. A written of
several gigabytes moments after boot means the rollback did not happen. The
snapshot listing should show @blank and, if Trap 4 applies to you,
nothing else — which is the point at which people discover they have been
destroying their auto-snapshots.
The test¶
test.nix in this directory is a NixOS VM test that runs all three checks
above against a real machine, across a real reboot:
or, from a flake:
checks.x86_64-linux.zfs-impermanence-rollback =
pkgs.callPackage inputs.recipes + "/modules/zfs-impermanence-rollback/test.nix" { };
It builds a real pool on a scratch disk, boots a specialisation whose root is
tank/root with a second neededForBoot dataset at /state and an unwiped
one at /persist, writes a marker into each, reboots, and asserts:
- the marker on the root dataset is gone;
- the marker on the second
neededForBootdataset is gone, and the dataset is empty — this is the Trap 1 assertion, and it fails on its own when only that dataset's rollback is disabled while the machine still boots cleanly; - the marker on
/persistsurvives, so the two above are not just "the disk came up blank"; - inside the initrd,
zfs-wipe-state.serviceisBefore=sysroot.mount sysroot-state.mount, both wipes returnedsuccess, and each wipe'sActiveEnterTimestampMonotonicprecedes its mount unit'sInactiveExitTimestampMonotonic— the ordering as observed at runtime, not merely as declared.
Plus four eval-time assertions that need no VM: enforceNeededForBoot really
sets it, the emitted unit really carries the sysroot.mount anchor and the
pool import ordering, and the neededForBoot guard rail really fires (and
fires only then).
Note the narrow-anchor case: if the module is changed to emit
Before=sysroot-state.mount alone, every behavioural assertion above still
passes — the machine wipes correctly and looks perfect. Only the recorded
Before= catches it. That is Trap 1 reproduced in miniature, and it is why
that assertion is in the test.
Caveats¶
- This module does not create the blank snapshot. By design: creating it from a module means creating it at some arbitrary later moment, from whatever state the dataset is in then. See Trap 5.
- It does not persist anything. Pair it with
nix-community/impermanence'senvironment.persistenceor your own bind mounts. Getting the wipe right and the persistence wrong is worse than not wiping. - Systemd initrd only. Asserted. Under scripted stage 1 there is no unit
ordering to be correct about, and
boot.initrd.postDeviceCommandsis removed anyway. mountPoint = nullis supported but weaker. The dataset is still wiped, still anchored onsysroot.mount, but theneededForBootguard rail and the device/dataset cross-check cannot run. Use it only for datasets that are not infileSystemsat all.- Nested datasets are not recursive. zfs-rollback(8): "The
-rRoptions do not recursively destroy the child snapshots of a recursive snapshot. Only direct snapshots of the specified filesystem are destroyed." Rolling backrpool/local/homedoes nothing to a child dataset such asrpool/local/home/<someone>. Declare each dataset you want wiped as its own entry. - Btrfs needs a different implementation. The equivalent (
btrfs subvolume delete+snapshotfrom a read-only blank, under the top-levelsubvolid=5mount) has the same ordering requirement and the same silent-failure shape, plus one extra: nested subvolumes created after the fact — bysystemd-nspawn, by container runtimes — must be deleted first or the parentdeletefails. This module is ZFS-only. - Version bounds. Written against nixpkgs
26.11(nixos/modules/tasks/ filesystems.nix,.../filesystems/zfs.nix,.../system/boot/systemd/ initrd.nix), systemd 261, OpenZFS 2.4. The unit-name derivation mirrorsgetPoolMounts; if upstream changes how initrd mount units are named, that is the one thing here that has to change with it — which is another argument for anchoring onsysroot.mount, whose name has been stable for the entire life of the systemd initrd.
License¶
CC0-1.0.
Source¶
modules/zfs-impermanence-rollback/default.nix
{
config,
lib,
utils,
...
}:
let
inherit (lib)
mkEnableOption
mkIf
mkOption
types
;
cfg = config.zfsWipeOnBoot;
# The initrd mount unit for a filesystem is the /sysroot-prefixed, systemd-
# escaped mount point. This mirrors `getPoolMounts` in
# nixos/modules/tasks/filesystems/zfs.nix, including the trailing-slash strip
# that keeps "/" from becoming "sysroot-.mount".
initrdMountUnit =
mountPoint: "${utils.escapeSystemdPath ("/sysroot" + (lib.removeSuffix "/" mountPoint))}.mount";
poolOf = dataset: lib.head (lib.splitString "/" dataset);
enabled = lib.filterAttrs (_: e: e.enable) cfg.datasets;
# sysroot.mount is the ONLY anchor that is always present and always ordered
# ahead of every other /sysroot/* mount (systemd.mount(5): a mount unit
# beneath another in the hierarchy gains an implicit Requires= and After= on
# the parent). Naming the dataset's own mount unit as well is redundant but
# self-documenting, and it costs nothing.
beforeUnits =
e:
lib.unique (
[ "sysroot.mount" ] ++ lib.optional (e.mountPoint != null) (initrdMountUnit e.mountPoint)
);
mkScript = e: ''
snap="${e.dataset}@${e.snapshot}"
if ! zfs list -H -t snapshot -o name "$snap" > /dev/null 2>&1; then
${
{
fail = ''
echo "zfs-wipe-on-boot: $snap does not exist. Refusing to boot with" >&2
echo "state that was supposed to be discarded. Create it with:" >&2
echo " zfs snapshot $snap" >&2
exit 1
'';
create = ''
echo "zfs-wipe-on-boot: $snap missing, creating it from the CURRENT" >&2
echo "contents of ${e.dataset}. This boot keeps its state." >&2
zfs snapshot "$snap"
'';
ignore = ''
echo "zfs-wipe-on-boot: $snap missing, skipping rollback" >&2
exit 0
'';
}
.${e.onMissingSnapshot}
}
fi
zfs rollback ${lib.optionalString e.recursive "-r "}"$snap"
'';
mkService = e: {
description = "Wipe ${e.dataset} by rolling back to @${e.snapshot}";
wantedBy = [ "initrd.target" ];
after = [ "zfs-import-${poolOf e.dataset}.service" ] ++ e.after;
requires = e.requires;
before = beforeUnits e;
unitConfig = {
DefaultDependencies = "no";
}
// lib.optionalAttrs e.failHard {
OnFailure = "emergency.target";
OnFailureJobMode = "replace-irreversibly";
};
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
path = [ cfg.package ];
script = mkScript e;
};
mounted = lib.filterAttrs (_: e: e.mountPoint != null) enabled;
# Which declared filesystem actually governs `path`? The DEEPEST matching
# mount point wins, so a persisted dataset nested under a wiped one (the
# usual /persist-under-/ layout) is correctly seen as safe.
governingMount =
path:
let
candidates = lib.filter (
mp: mp == "/" || mp == path || lib.hasPrefix (mp + "/") path
) (lib.attrNames config.fileSystems);
in
lib.foldl' (a: b: if lib.stringLength b > lib.stringLength a then b else a) "/" candidates;
wipedMountPoints = lib.mapAttrsToList (_: e: e.mountPoint) mounted;
# A path is destroyed on every boot iff the mount governing it is wiped.
onWipedPath = path: lib.elem (governingMount path) wipedMountPoints;
# Only reason about real filesystem paths. Store paths are immutable and
# irrelevant here; a non-absolute value is not ours to interpret.
plainPath = p: let s = toString p; in lib.hasPrefix "/" s && !(lib.hasPrefix "/nix/store/" s);
atRisk = paths: lib.filter (p: plainPath p && onWipedPath (toString p)) paths;
sshdHostKeyPaths = lib.optionals (config.services.openssh.enable or false) (
map (k: k.path) (config.services.openssh.hostKeys or [ ])
);
initrdHostKeyPaths = lib.optionals (config.boot.initrd.network.ssh.enable or false) (
config.boot.initrd.network.ssh.hostKeys or [ ]
);
fsOf = e: config.fileSystems.${e.mountPoint} or null;
datasetModule =
{ ... }:
{
options = {
enable = mkOption {
type = types.bool;
default = true;
description = "Whether to wipe this dataset on every boot.";
};
dataset = mkOption {
type = types.str;
example = "rpool/local/home";
# No default: guessing a dataset name from an attribute key is how a
# wipe ends up pointed at nothing.
description = ''
Full ZFS dataset name. The pool component is used to derive the
initrd import unit this rollback is ordered after
(`zfs-import-<pool>.service`).
'';
};
snapshot = mkOption {
type = types.str;
default = cfg.snapshot;
defaultText = lib.literalExpression "config.zfsWipeOnBoot.snapshot";
description = ''
Snapshot name (without the `@`) representing the empty, known-good
state. It must already exist before the first boot — create it at
install time, e.g. from a disko `postCreateHook`.
'';
};
mountPoint = mkOption {
type = types.nullOr types.path;
default = null;
example = "/home";
description = ''
Where this dataset is mounted. Setting it turns on the
`neededForBoot` guard rail (see `enforceNeededForBoot`) and the
matching assertions. `null` means the dataset is not one of this
host's `fileSystems`, in which case the only ordering anchor is
`sysroot.mount`.
'';
};
recursive = mkOption {
type = types.bool;
default = true;
description = ''
Pass `-r` to `zfs rollback`, destroying any snapshot and bookmark
newer than the target. Required for an unattended wipe: without it
the rollback FAILS the moment anything (an auto-snapshot timer, a
replication job) has taken a newer snapshot.
The flip side: on a dataset with `com.sun:auto-snapshot=true` this
silently destroys every automatic snapshot of that dataset on every
boot. Do not enable auto-snapshots on a wiped dataset.
'';
};
onMissingSnapshot = mkOption {
type = types.enum [
"fail"
"create"
"ignore"
];
default = "fail";
description = ''
What to do when the blank snapshot does not exist.
- `fail` — abort. Combined with `failHard` this stops the boot
instead of silently keeping state.
- `create` — snapshot the dataset's *current* contents under that
name and continue. Convenient when adding a dataset to an
existing host; this boot keeps whatever was there.
- `ignore` — skip the rollback and continue. Only ever correct for
a dataset you are in the middle of decommissioning.
'';
};
failHard = mkOption {
type = types.bool;
default = true;
description = ''
Add `OnFailure=emergency.target` /
`OnFailureJobMode=replace-irreversibly` so a failed wipe drops to
the initrd emergency shell rather than booting with state that was
supposed to be gone.
Set to `false` on remote hosts that have neither
`boot.initrd.systemd.emergencyAccess` nor initrd SSH: there, the
emergency target is an unreachable brick.
'';
};
after = mkOption {
type = types.listOf types.str;
default = [ ];
example = [ "systemd-cryptsetup@crypted.service" ];
description = ''
Extra `After=` units. The pool's initrd import unit is added
automatically; this is for anything the import itself depends on
that is not already wired, such as a LUKS container.
'';
};
requires = mkOption {
type = types.listOf types.str;
default = [ ];
description = "Extra `Requires=` units.";
};
};
};
in
{
options.zfsWipeOnBoot = {
enable = mkEnableOption "wiping ZFS datasets to a blank snapshot in the initrd";
package = mkOption {
type = types.package;
default = config.boot.zfs.package;
defaultText = lib.literalExpression "config.boot.zfs.package";
description = ''
ZFS package providing `zfs`. Must be the same build the initrd imports
the pool with, or the rollback can hit a feature-flag mismatch.
'';
};
snapshot = mkOption {
type = types.str;
default = "blank";
description = "Default snapshot name for every entry in `datasets`.";
};
namePrefix = mkOption {
type = types.str;
default = "zfs-wipe";
description = ''
Prefix for the generated initrd unit names: `<prefix>-<key>.service`.
'';
};
enforceNeededForBoot = mkOption {
type = types.bool;
default = true;
description = ''
Set `fileSystems.<mountPoint>.neededForBoot = true` for every wiped
dataset that names a mount point.
This is not a convenience. `neededForBoot` is what puts the filesystem
into the initrd fstab, and therefore what makes its
`sysroot-*.mount` unit exist at all. Without it the mount happens in
stage 2, long after the initrd has been torn down, and the wipe races
against nothing that systemd can order it against.
'';
};
datasets = mkOption {
type = types.attrsOf (types.submodule datasetModule);
default = { };
example = lib.literalExpression ''
{
root = { dataset = "rpool/local/root"; mountPoint = "/"; };
home = { dataset = "rpool/local/home"; mountPoint = "/home"; };
}
'';
description = ''
Datasets to roll back to their blank snapshot on every boot. The
attribute name is used for the unit name and, unless overridden, as
the dataset name.
'';
};
};
config = mkIf (cfg.enable && enabled != { }) {
assertions = [
{
assertion = config.boot.initrd.systemd.enable;
message = ''
zfsWipeOnBoot requires the systemd initrd
(`boot.initrd.systemd.enable = true`). Scripted stage 1 has no unit
ordering at all: `boot.initrd.postDeviceCommands` runs as an
unordered blob and is a removed option under systemd stage 1
(nixos/modules/system/boot/systemd/initrd.nix, the `obsoleteOpt`
list).
'';
}
]
++ lib.mapAttrsToList (name: e: {
assertion = fsOf e != null;
message = ''
zfsWipeOnBoot.datasets.${name}.mountPoint is "${toString e.mountPoint}"
but there is no `fileSystems."${toString e.mountPoint}"`. Either declare
the filesystem or set mountPoint = null.
'';
}) mounted
++ lib.mapAttrsToList (name: e: {
assertion = fsOf e == null || utils.fsNeededForBoot (fsOf e);
message = ''
zfsWipeOnBoot.datasets.${name} wipes ${e.dataset} mounted at
${toString e.mountPoint}, but that filesystem is not neededForBoot.
Its `${initrdMountUnit e.mountPoint}` unit therefore does not exist in
the initrd, the mount happens in stage 2 instead, and the wipe is
unordered with respect to it. Set `neededForBoot = true` (or leave
zfsWipeOnBoot.enforceNeededForBoot at its default).
'';
}) mounted
++ [
{
assertion = atRisk sshdHostKeyPaths == [ ];
message = ''
zfsWipeOnBoot would destroy this host's SSH host key on every boot:
${lib.concatStringsSep ", " (map toString (atRisk sshdHostKeyPaths))}
services.openssh.hostKeys must live on a dataset that is NOT rolled
back, or every reboot regenerates the machine's identity. sshd still
starts, so nothing looks broken -- but every client that pinned the
old key refuses to connect, and you find out from the one machine you
can no longer reach.
'';
}
{
assertion = atRisk initrdHostKeyPaths == [ ];
message = ''
zfsWipeOnBoot would destroy the INITRD SSH host key:
${lib.concatStringsSep ", " (map toString (atRisk initrdHostKeyPaths))}
boot.initrd.network.ssh.hostKeys is read from the live filesystem when
the initrd secrets are assembled, so a wiped path loses it. On a
remote-unlock host this is the worst version of the failure: the key
that changed is the one used by whoever has to log in to unlock the
disk, and there is no other way in.
'';
}
]
++ lib.mapAttrsToList (name: e: {
assertion = fsOf e == null || (fsOf e).fsType == "zfs";
message = ''
zfsWipeOnBoot.datasets.${name} wipes the ZFS dataset ${e.dataset}, but
fileSystems."${toString e.mountPoint}".fsType is
"${(fsOf e).fsType}". Rolling a dataset back underneath a mount of a
different filesystem does nothing useful.
'';
}) mounted;
warnings = lib.filter (w: w != null) (
lib.mapAttrsToList (
name: e:
if fsOf e != null && (fsOf e).device != null && (fsOf e).device != e.dataset then
''
zfsWipeOnBoot.datasets.${name} rolls back "${e.dataset}" but
fileSystems."${toString e.mountPoint}".device is
"${(fsOf e).device}". If those are not the same dataset, the wipe
is happening somewhere nobody is looking.
''
else
null
) mounted
);
fileSystems = mkIf cfg.enforceNeededForBoot (
lib.mapAttrs' (_: e: lib.nameValuePair e.mountPoint { neededForBoot = true; }) mounted
);
boot.initrd.systemd.services = lib.mapAttrs' (
name: e: lib.nameValuePair "${cfg.namePrefix}-${name}" (mkService e)
) enabled;
};
}