systemd-boot-mirrored-esp¶
Modules
Two EFI system partitions on two different disks, kept byte-identical, plus a pinned recovery boot entry that the bootloader's own garbage collector cannot reap — and the ZFS mount barrier that makes "is the data actually there?" a real precondition instead of a hope.
Three separate mechanisms, one module, because they fail as a unit: a mirror of a broken ESP is a broken ESP, a recovery entry whose kernel was deleted is worse than no recovery entry, and a recovery payload staged on an unmounted dataset is silently not there.
The problem¶
1. systemd-boot has no mirroring at all¶
nixpkgs' GRUB module has boot.loader.grub.mirroredBoots
(nixos/modules/system/boot/loader/grub/grub.nix:276) — a list of
{ path, devices } pairs, and the installer walks it and installs to every one.
The systemd-boot module has no equivalent. It writes exactly one $BOOT:
bootMountPoint =
if cfg.xbootldrMountPoint != null then cfg.xbootldrMountPoint else efi.efiSysMountPoint;
(nixos/modules/system/boot/loader/systemd-boot/systemd-boot.nix:65.) One
mount point, one partition. If that NVMe dies, the machine does not boot, no
matter how well mirrored the root pool is. Mirroring the pool and forgetting
the ESP is the single most common gap in a "redundant" two-disk NixOS box:
zroot survives, and the firmware has nothing to load.
2. configurationLimit deletes the recovery kernel¶
This is the trap that gives the recipe its reason to exist, and it is worth reading the upstream code rather than trusting a summary.
boot.loader.systemd-boot.configurationLimit is documented as "maximum number
of latest generations in the boot menu"
(nixos/modules/system/boot/loader/systemd-boot/systemd-boot.nix:234). It
sounds like a display limit. It is not. It is fed straight into the installer:
(systemd-boot-builder.py:33 and :455.) Everything downstream — the keep-set
for the ESP — is computed from that truncated list. Then:
def garbage_collect(gc_roots: BootFileList) -> None:
keep = {BOOT_MOUNT_POINT / gc_root.path for gc_root in gc_roots}
def delete_path(e: os.DirEntry) -> None:
if e.is_file(follow_symlinks=True) and Path(e.path) not in keep:
os.remove(e.path)
for e in os.scandir(BOOT_MOUNT_POINT / NIXOS_DIR):
delete_path(e)
(systemd-boot-builder.py:623.) It scandirs $BOOT/EFI/nixos and deletes
every file there that is not owned by one of the last N generations. Not
"files it wrote". Not "files matching a pattern". Every file.
So the obvious way to build a recovery entry — boot into a known-good
generation, note its kernel and initrd under EFI/nixos/, hand-write a
.conf pointing at them — produces an entry that works today, works after the
next deploy, and is dangling by the sixth deploy, with
configurationLimit = 5. The failure mode is precisely inverted from what you
want: the recovery entry is present in the menu right up until you have done
enough deploys to actually need it, and then selecting it drops you at
Error: not found from the stub loader, on a machine you cannot log into.
Nothing warns. Nothing fails. The deploy that deletes the recovery kernel is green.
3. And nix-collect-garbage deletes the source¶
The reflex fix is boot.loader.systemd-boot.extraFiles, which is re-copied
after garbage collection (see Trap 3 below for the exact ordering). It does not
solve this problem, because its type is types.attrsOf types.path
(systemd-boot.nix:392) — it can only pin something Nix can evaluate to a store
path. The kernel of a past generation is a store path you would have to write
literally into the config, and a literal /nix/store/... string carries no
string context, so nothing roots it and nix-collect-garbage -d deletes it. The
next deploy then fails in install -Dp with No such file or directory, in the
middle of bootloader installation, on a system whose ESP has already been
garbage collected.
The recovery payload has to survive two independent garbage collectors — the ESP's and the Nix store's — plus a store wipe and a reinstall. The only place that satisfies all of those is a plain directory on persistent storage, outside the store and outside the ESP.
Traps¶
Trap 1 — the recovery entry must NOT be named nixos-*.conf¶
The entry-side garbage collection is regex-driven:
for e in os.scandir(BOOT_MOUNT_POINT / "loader" / "entries"):
match = re.fullmatch(r"nixos-.+\.conf", e.name)
if match:
delete_path(e)
(systemd-boot-builder.py:633-636.) Entries that do not match are left
alone — that is what makes a hand-placed recovery entry survivable at all. But
the natural name for a recovery entry is something like
nixos-recovery.conf, which matches nixos-.+\.conf exactly, is not in the
keep set, and is therefore deleted on every single install.
This module hard-asserts against it:
Name it recovery-shell-init.conf, zz-recovery.conf, anything that does not
start with nixos-. Use the sort-key field inside the entry to control where
it lands in the menu; the filename is not the ordering mechanism.
Trap 2 — extraInstallCommands is the only hook that runs late enough¶
The installer script is assembled as:
finalSystemdBootBuilder = pkgs.writeScript "install-systemd-boot.sh" ''
#!${pkgs.runtimeShell}
set -euo pipefail
${systemdBootBuilder}/bin/systemd-boot "$@"
${cfg.extraInstallCommands}
'';
(systemd-boot.nix:104-109.) extraInstallCommands runs after the Python
builder has completely finished. Inside that builder the order is:
garbage_collect(boot_files)— pruneEFI/nixosandnixos-*.confwrite_boot_files(...)— install the live generationswrite_loader_conf(...)remove_extra_files()— delete the previous run'sextraFiles/extraEntriesrun([COPY_EXTRA_FILES])— re-copy them
(systemd-boot-builder.py:589-596.) Anything that wants to place a file on the
ESP and have it stay must run after step 5. extraInstallCommands is the only
NixOS-level hook that does. This module emits exactly one block there:
if [ -d /var/lib/boot-recovery ]; then
.../cp -f \
/var/lib/boot-recovery/<kernel>.efi \
/var/lib/boot-recovery/<initrd>.efi \
/boot/EFI/nixos/
.../cp -f \
/var/lib/boot-recovery/recovery-shell-init.conf \
/boot/loader/entries/recovery-shell-init.conf
fi
.../rsync -a --delete /boot/ /boot-mirror/
The restore is unconditional-per-deploy, not one-shot: it re-lays the files after every garbage collection, so the pin is re-established as fast as it is broken.
Trap 3 — the mirror must be the LAST statement, and it must --delete¶
Order matters twice over:
- The
rsyncruns after the recovery restore. Reverse the two and the mirror ESP is a snapshot taken one instant before the recovery entry is put back — permanently missing exactly the entry you built the second disk for. --deleteis not optional. Without it the mirror is a union of every generation ever installed: it grows monotonically, fills a 1 GiB FAT partition, and from then on rsync fails mid-copy and the mirror silently diverges from the primary. A mirror you cannot trust is worse than none, because you will try to boot it.
The trailing slashes are load-bearing: rsync -a src/ dst/ copies the
contents of src; rsync -a src dst/ creates dst/src. The module always
emits both slashes.
Trap 4 — mountpoint -q, never test -d¶
The ZFS barrier polls with mountpoint -q, which reads the mount table. The
obvious alternative, test -d /data, is worse than useless: when a pool fails
to import, the mountpoint directory still exists on the underlying root
filesystem, empty. test -d succeeds, the barrier reports green, and every
consumer starts writing into the root filesystem at a path that is supposed to
be a 40 TB pool.
Then the pool imports late and ZFS mounts over the top. ZFS's overlay
property defaults to on, so the mount does not fail with "directory not
empty" — it succeeds and silently shadows everything that was written
underneath. The data is not lost, exactly; it is invisible, on the wrong
filesystem, filling the root pool, and it reappears the next time the data pool
fails to import. Nobody finds this quickly.
Trap 5 — after AND requires, both, on the import unit¶
Both, always. requires alone pulls the import service into the transaction but
imposes no ordering, so the barrier can run first and burn all its retries while
the import has not started. after alone orders correctly but only if the
import is in the same transaction — if nothing pulled it in, After= on an
inactive unit is a no-op and the barrier is ordered after nothing.
The barrier is also wantedBy = [ "multi-user.target" ] by default rather than
being pulled in solely by its consumers. A barrier that only runs when someone
needs it is a barrier whose failure is invisible on a host where that consumer
happens to be disabled.
Trap 6 — there is no .mount unit to order against¶
The reflex is RequiresMountsFor=/data, or after = [ "data.mount" ]. Neither
exists for a natively-mounted ZFS dataset.
Datasets with mountpoint=legacy get a real fileSystems entry and therefore a
real systemd .mount unit. Datasets with a native mountpoint=/data are
mounted by zfs-mount.service
(nixos/modules/tasks/filesystems/zfs.nix:945), a single oneshot that runs
zfs mount -a for everything at once. systemd has no per-dataset unit to bind
to, and zfs-mount.service reports success even when an individual dataset
failed to mount.
So there is nothing to order against, and polling is not laziness — it is the only mechanism that actually observes the property you care about. The backoff is bounded (7 attempts, 1+2+4+…+64 = 127 s) so a genuinely dead pool fails the unit instead of hanging boot forever.
Trap 7 — the recovery restore fails OPEN, on purpose¶
if [ -d <directory> ] means: no payload staged, no restore, deploy proceeds.
That is deliberate. A host must be deployable before its recovery payload has
been staged — otherwise the very first deploy of a new machine is blocked on a
chicken-and-egg problem. But it also means a disappeared payload is silent,
and the most likely way for it to disappear is that it lives on a dataset that
did not mount. [ -d /data/boot-recovery ] is false when /data is an empty
unmounted directory, and the deploy is green.
If you stage the payload on a ZFS dataset, put a barrier on that dataset and order something you actually watch behind it. The two halves of this module are in the same file for that reason.
Trap 8 — identify both ESPs by filesystem UUID¶
primary.device and mirror.device are applied with mkForce, overriding
whatever disko generated, and the README example uses /dev/disk/by-uuid/.
by-label/by-partlabel: a disko layout that names both partitionsESPgives two partitions with the same label. Which one/bootresolves to is then a race, and the loser getsrsync --deleted onto the winner.by-id/by-path: encodes the controller slot. Move a disk after replacing the dead one and the mount points swap.by-uuid:rsynccopies files, not the filesystem. The FAT volume ID lives in the boot sector and is never touched by the mirroring, so the two UUIDs stay distinct for the life of the partitions — including after a failover.
Trap 9 — the mirror is not in NVRAM, but it is on the fallback path¶
bootctl install writes an EFI boot variable pointing at the primary ESP's
partition. The mirror is never registered: it is a passive replica that
systemd-boot has never heard of.
What makes it bootable anyway is that bootctl install also writes the
removable-media fallback, EFI/BOOT/BOOTX64.EFI — and rsync -a --delete
copies that to the mirror along with everything else. Most firmwares will boot
the second disk from the fallback path once the first is gone or deselected, and
all of them let you pick it manually from the firmware boot menu.
Test this before you need it, by disabling the primary disk in firmware setup and booting. A mirror nobody has ever booted is a hypothesis.
Trap 10 — loader/random-seed is per-installation state¶
rsyncExcludes defaults to [ ] — a byte-for-byte replica, which is what makes
the mirror boot with zero fixups.
The cost is that loader/random-seed is duplicated. bootctl treats that file
as per-installation state and warns against carrying it into a cloned image; the
seed is credited to the kernel entropy pool at boot and refreshed afterwards, so
two copies means the same seed can be credited twice — once from each ESP, if
you ever boot both. On a single machine with one live ESP at a time the exposure
is small, but if you would rather have systemd-boot regenerate it on first boot
of the mirror:
Decide once and write it down; the default is stated here so it is a choice rather than an accident.
Trap 11 — size the ESP for configurationLimit, then double the disks¶
A NixOS generation costs roughly 15 MB of kernel plus 60–120 MB of initrd on the
ESP, more with boot.initrd.includeDefaultModules and firmware blobs. At
configurationLimit = 5 that is ~400–700 MB, plus the pinned recovery pair,
plus systemd-boot itself. A 512 MB ESP overflows; 1–1.5 GiB is comfortable.
An overflowing ESP is a nasty failure because it happens inside
write_boot_files, after garbage_collect has already run — the old
generations are gone and the new one did not fit.
Usage¶
{
imports = [ ./systemd-boot-mirrored-esp ];
boot.loader.systemd-boot = {
enable = true;
configurationLimit = 5;
};
boot.mirroredEsp = {
enable = true;
primary = {
mountPoint = "/boot";
device = "/dev/disk/by-uuid/1234-ABCD";
};
mirror = {
mountPoint = "/boot-mirror";
device = "/dev/disk/by-uuid/5678-EF01";
};
recovery = {
enable = true;
directory = "/persistent/boot-recovery";
efiFiles = [
"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa-linux-6.12.0-bzImage.efi"
"bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb-initrd-linux-6.12.0-initrd.efi"
];
entryFiles = [ "recovery-shell-init.conf" ];
};
};
boot.zfsMountBarriers = {
data = {
mountpoint = "/data";
importUnits = [ "zfs-import-datapool.service" ];
};
data-archive = {
mountpoint = "/data/archive";
importUnits = [ "zfs-import-datapool.service" ];
};
};
}
A consumer then orders itself behind a barrier:
systemd.services.my-archiver = {
after = [ "wait-for-zfs-data-archive.service" ];
requires = [ "wait-for-zfs-data-archive.service" ];
};
Staging the recovery payload (one-time, imperative)¶
The payload is deliberately not declarative — that is the whole point. Boot the generation you want to be able to fall back to, then, as root:
mkdir -p /persistent/boot-recovery
cp /boot/EFI/nixos/*-linux-*-bzImage.efi /persistent/boot-recovery/
cp /boot/EFI/nixos/*-initrd-linux-*.efi /persistent/boot-recovery/
cp /boot/loader/entries/nixos-generation-<N>.conf \
/persistent/boot-recovery/recovery-shell-init.conf
Then edit the copied .conf: rename the title, add a sort-key so it lands
where you want in the menu, and consider appending init=… overrides such as
systemd.unit=rescue.target. Leave the linux and initrd lines pointing at
/EFI/nixos/<the exact filenames you copied> — those filenames go into
recovery.efiFiles verbatim.
The options init=/nix/store/…-nixos-system-…/init line is the part that keeps
this honest: it names a store path. Keep that generation pinned as a Nix GC root
(nix-env -p /nix/var/nix/profiles/system --list-generations, then do not
--delete-older-than past it) or accept that the recovery entry boots a kernel
and initrd into an emergency shell without a working init. For a
"get me a shell on this box" recovery entry the latter is often enough — the
initrd is self-contained — but know which one you built.
Verify after every deploy that changes the kernel:
Options¶
| Option | Default | Effect |
|---|---|---|
boot.mirroredEsp.enable |
false |
Nothing applies until this is on. |
boot.mirroredEsp.primary.mountPoint |
"/boot" |
Must equal the partition systemd-boot writes to (asserted). |
boot.mirroredEsp.primary.device |
null |
mkForced into fileSystems. Use by-uuid — Trap 8. |
boot.mirroredEsp.mirror.enable |
true |
Emit the mirroring rsync. |
boot.mirroredEsp.mirror.mountPoint |
"/boot-mirror" |
Passive replica; must differ from primary (asserted). |
boot.mirroredEsp.mirror.device |
null |
mkForced into fileSystems. |
boot.mirroredEsp.mountOptions |
[ "umask=0077" ] |
Added to both ESPs. Concatenates with disko's defaults. |
boot.mirroredEsp.rsyncFlags |
[ "-a" "--delete" ] |
See Trap 3 before changing. |
boot.mirroredEsp.rsyncExcludes |
[ ] |
--exclude= list. See Trap 10. |
boot.mirroredEsp.recovery.enable |
false |
Restore the pinned payload after every install. |
boot.mirroredEsp.recovery.directory |
"/var/lib/boot-recovery" |
Must be outside the store and outside the ESP (asserted). |
boot.mirroredEsp.recovery.efiFiles |
[ ] |
Basenames copied into <ESP>/EFI/nixos/. |
boot.mirroredEsp.recovery.entryFiles |
[ ] |
Basenames copied into <ESP>/loader/entries/. Rejected if nixos-*.conf — Trap 1. |
boot.mirroredEsp.coreutilsPackage |
pkgs.coreutils |
Provides cp. |
boot.mirroredEsp.rsyncPackage |
pkgs.rsync |
Provides rsync. |
boot.mirroredEsp.utilLinuxPackage |
pkgs.util-linux |
Provides mountpoint for the barriers. |
boot.zfsMountBarriers.<name>.mountpoint |
— | Path polled with mountpoint -q. |
boot.zfsMountBarriers.<name>.importUnits |
— | Placed in both after and requires — Trap 5. |
boot.zfsMountBarriers.<name>.attempts |
7 |
Exponential backoff, 127 s total. |
boot.zfsMountBarriers.<name>.wantedBy |
[ "multi-user.target" ] |
Keep it, so failures are visible. |
Barriers are independent of boot.mirroredEsp.enable; a host can use either
half alone.
Testing¶
test.nix is a real NixOS VM test. Run it standalone, no flake
needed:
or from a flake, pkgs.callPackage ./modules/systemd-boot-mirrored-esp/test.nix { }.
The VM boots UEFI via virtualisation.useBootLoader, gets a second 512 MiB
virtio disk formatted as the mirror ESP and a third one for the pinned payload,
and then rolls through four generations with configurationLimit = 2 so the
installer's garbage collector really fires.
What it proves:
- The two ESPs are byte-identical. Not "both non-empty": the same recursive
file and directory listing, the same
sha256sumfor every regular file, a cleandiff -r, and a floor on file count and total size so the comparison cannot pass vacuously on an empty ESP. A stray file planted in the mirror beforehand is gone afterwards, which is what proves--delete(Trap 3). - The recovery kernel, initrd and loader entry are present after every install, on both ESPs, byte-identical to what was staged.
- The recovery entry survives generation garbage collection. Four
generations with three distinct initrds are built; the test asserts the
initrds really are distinct (otherwise the whole check would be vacuous),
that generations 1 and 2 lose their loader entries, that generation 2's
initrd file is deleted from
EFI/nixos— and that through all of it the pinned payload is still there and the entry'slinux/initrdlines still resolve to existing, correct files. - It actually boots. After the last install,
loader.confis pointed at the recovery entry, the VM is rebooted, and the test asserts a unique marker from the pinned entry'soptionsline shows up in/proc/cmdline. That string exists nowhere else, so it can only have come from firmware loading the pinned kernel and initrd off the ESP.
The instrument that keeps the survival assertions honest: before the first
install the test plants two unpinned decoys — one file in EFI/nixos and one
nixos-*.conf in loader/entries — and asserts that both are gone after
the install. If they were still there, "the recovery files survived" would only
mean "the collector never ran", and the test says so instead of passing.
Two things the module declares cannot be observed from inside a VM at all,
because qemu-vm.nix replaces fileSystems wholesale with
mkVMOverride config.virtualisation.fileSystems: the mkForced ESP devices
and the umask=0077 mount options. Those are checked in a plain non-VM
evaluation at the top of test.nix, together with the two assertions from
Trap 1 (nixos-*.conf entry name) and the payload-inside-the-ESP guard. The VM
then mounts both ESPs by hand with exactly the options that evaluation proved
the module declares.
Not covered: firmware-level failover (pulling the primary disk and booting the
mirror ESP), which needs a second NVRAM boot entry and disk removal the test
driver cannot express; Secure Boot signing of the mirrored/pinned binaries; and
xbootldrMountPoint layouts, which the module does not support for mirroring
anyway.
Caveats¶
- Nothing verifies the recovery entry. The module guarantees the files are
re-laid after every install and that their names cannot be garbage collected.
It cannot check that the
.confpoints at kernels that exist, or that they boot. Boot it once a quarter. - The mirror is refreshed only when the bootloader is installed. That is
every
nixos-rebuild switch/boot, but not on anixos-rebuild test, and not if you hand-edit something on the primary ESP.diff -ris the check. - XBOOTLDR is not supported for mirroring. With
boot.loader.systemd-boot.xbootldrMountPointset, entries and kernels live on XBOOTLDR while the EFI binaries live on the ESP, so a single-partition mirror is not a bootable copy. The assertion forcesprimary.mountPointto the partition that receives entries; mirroring the other one is out of scope. fileSystems.<p>.optionsconcatenates.mountOptionsis added to whatever else declared the mount; it does not replace it. If disko already contributesdefaults, the result is[ "defaults" "umask=0077" ].- Secure Boot changes the picture. With a shim/
sbctlsetup the mirrored binaries must be signed with the same keys, and a hand-staged recovery kernel is unsigned unless you signed it before staging. Sign first, stage second. - This is not a substitute for a rescue USB. It covers "one disk died" and
"the last five generations are all broken". It does not cover a corrupted
root pool, and it never will — see
remote-luks-unlockfor getting into a box whose initrd is waiting for a passphrase, andzfs-native-encryption-keysfor the key handling that has to work before any of this matters.
Source¶
modules/systemd-boot-mirrored-esp/default.nix
{
lib,
pkgs,
config,
...
}:
let
inherit (lib)
concatMap
concatStringsSep
mkEnableOption
mkForce
mkIf
mkMerge
mkOption
mkPackageOption
optionalString
optionals
types
;
cfg = config.boot.mirroredEsp;
sdb = config.boot.loader.systemd-boot;
esp = cfg.primary.mountPoint;
cp = "${cfg.coreutilsPackage}/bin/cp";
rsync = "${cfg.rsyncPackage}/bin/rsync";
isNixosEntryName = n: lib.hasPrefix "nixos-" n && lib.hasSuffix ".conf" n;
badEntryNames = builtins.filter isNixosEntryName cfg.recovery.entryFiles;
dir = cfg.recovery.directory;
kernelBlock = optionals (cfg.recovery.efiFiles != [ ]) (
[ " ${cp} -f \\" ]
++ map (f: " ${dir}/${f} \\") cfg.recovery.efiFiles
++ [ " ${esp}/EFI/nixos/" ]
);
entryBlocks = concatMap (e: [
" ${cp} -f \\"
" ${dir}/${e} \\"
" ${esp}/loader/entries/${e}"
]) cfg.recovery.entryFiles;
recoveryLines = optionals cfg.recovery.enable (
[ "if [ -d ${dir} ]; then" ] ++ kernelBlock ++ entryBlocks ++ [ "fi" ]
);
rsyncArgs = cfg.rsyncFlags ++ map (p: "--exclude=${p}") cfg.rsyncExcludes;
mirrorLines = optionals cfg.mirror.enable [
"${rsync} ${concatStringsSep " " rsyncArgs} ${esp}/ ${cfg.mirror.mountPoint}/"
];
installLines = recoveryLines ++ mirrorLines;
installScript = optionalString (installLines != [ ]) (concatStringsSep "\n" installLines + "\n");
barrierUnit = barrier: {
description = "Wait for ZFS native mount at ${barrier.mountpoint}";
after = barrier.importUnits;
requires = barrier.importUnits;
inherit (barrier) wantedBy;
serviceConfig = {
Type = "oneshot";
RemainAfterExit = true;
};
path = [ cfg.utilLinuxPackage ];
script = ''
attempts=0
delay=1
while [ $attempts -lt ${toString barrier.attempts} ]; do
if mountpoint -q "${barrier.mountpoint}"; then
echo "${barrier.mountpoint} is mounted"
exit 0
fi
attempts=$((attempts + 1))
echo "Waiting for ${barrier.mountpoint} (attempt $attempts/${toString barrier.attempts}, sleeping ''${delay}s)"
sleep $delay
delay=$((delay * 2))
done
echo "ERROR: ${barrier.mountpoint} not mounted after ${toString barrier.attempts} attempts"
exit 1
'';
};
barrierType = types.submodule {
options = {
mountpoint = mkOption {
type = types.str;
example = "/tank";
description = ''
Absolute path that must be a real mountpoint before the barrier unit
reports success. Checked with `mountpoint -q`, which tests the mount
table — not `test -d`, which succeeds on the empty directory left
behind by a pool that failed to import.
'';
};
importUnits = mkOption {
type = types.listOf types.str;
example = [ "zfs-import-tank.service" ];
description = ''
Units placed in BOTH `after` and `requires`. `requires` alone pulls
the import in but does not order against it; `after` alone orders but
lets the barrier run in a transaction where the import was never
queued. Both are needed.
'';
};
attempts = mkOption {
type = types.ints.positive;
default = 7;
description = ''
Poll attempts with exponential backoff (1s, 2s, 4s, ...). 7 attempts
is 127s of total sleep, which covers a slow spinning-disk pool import
without letting a genuinely broken pool hang the boot forever.
'';
};
wantedBy = mkOption {
type = types.listOf types.str;
default = [ "multi-user.target" ];
description = ''
Where the barrier is pulled in from. Keep `multi-user.target` so the
barrier runs even when no consumer happens to be enabled — otherwise
the failure is invisible until something else breaks.
'';
};
};
};
in
{
options.boot.mirroredEsp = {
enable = mkEnableOption "a mirrored EFI system partition with a pinned systemd-boot recovery entry";
primary = {
mountPoint = mkOption {
type = types.str;
default = "/boot";
description = ''
Mount point of the ESP that systemd-boot actually installs into.
Must equal `boot.loader.efi.efiSysMountPoint` (or
`boot.loader.systemd-boot.xbootldrMountPoint` when that is set) —
asserted below.
'';
};
device = mkOption {
type = types.nullOr types.str;
default = null;
example = "/dev/disk/by-uuid/1234-ABCD";
description = ''
Device for the primary ESP, applied with `mkForce` so it wins over a
disko-generated `fileSystems` entry. Use a filesystem UUID: the two
ESPs are byte-identical replicas, so `by-label` and `by-partlabel`
are ambiguous and `by-id`/`by-path` change when a disk is moved to
another slot. `null` leaves whatever else declared the mount alone.
'';
};
};
mirror = {
enable = mkOption {
type = types.bool;
default = true;
description = "Mirror the primary ESP onto a second ESP after every bootloader install.";
};
mountPoint = mkOption {
type = types.str;
default = "/boot-mirror";
description = ''
Mount point of the second ESP. It is a passive replica: systemd-boot
never writes here, the rsync at the end of the install does.
'';
};
device = mkOption {
type = types.nullOr types.str;
default = null;
example = "/dev/disk/by-uuid/5678-EF01";
description = "Device for the mirror ESP. See `primary.device`.";
};
};
mountOptions = mkOption {
type = types.listOf types.str;
default = [ "umask=0077" ];
description = ''
Mount options added to BOTH ESPs. `umask=0077` keeps the vfat tree
root-only; without it a FAT ESP is world-readable and every unprivileged
process can read the initrd, which on hosts using `boot.initrd.secrets`
contains secrets.
These are *added* to whatever else defines the mount (disko usually
contributes `defaults`), because `fileSystems.<p>.options` is a
`listOf` and concatenates definitions.
'';
};
rsyncFlags = mkOption {
type = types.listOf types.str;
default = [
"-a"
"--delete"
];
description = ''
Flags for the mirroring rsync. `--delete` is the point: the mirror must
be an exact replica, not a union of every generation that was ever
installed, or it silently fills up and then diverges.
'';
};
rsyncExcludes = mkOption {
type = types.listOf types.str;
default = [ ];
example = [ "loader/random-seed" ];
description = ''
Paths passed as `--exclude=`. See the README on
`loader/random-seed`: `bootctl` treats it as per-installation state and
explicitly says a cloned image must not carry someone else's seed.
'';
};
recovery = {
enable = mkEnableOption "restoring a pinned recovery kernel/initrd/entry onto the ESP after every install";
directory = mkOption {
type = types.str;
default = "/var/lib/boot-recovery";
example = "/persist/boot-recovery";
description = ''
Directory holding the pinned recovery payload. It must live OUTSIDE
the Nix store and OUTSIDE the ESP:
- outside the store, because `nix-collect-garbage` reaps the old
generation's kernel and initrd;
- outside the ESP, because systemd-boot's installer deletes every
file under `<ESP>/EFI/nixos` that does not belong to one of the
last `configurationLimit` generations.
The install script is a no-op if the directory does not exist, so a
host can be deployed before the payload is staged.
'';
};
efiFiles = mkOption {
type = types.listOf types.str;
default = [ ];
example = [
"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa-linux-6.12.0-bzImage.efi"
"bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb-initrd-linux-6.12.0-initrd.efi"
];
description = ''
Basenames inside `recovery.directory` copied back into
`<ESP>/EFI/nixos/` on every install. Use the exact names the
installer originally wrote, because the recovery `.conf` references
them by path.
'';
};
entryFiles = mkOption {
type = types.listOf types.str;
default = [ ];
example = [ "recovery-shell-init.conf" ];
description = ''
Basenames inside `recovery.directory` copied back into
`<ESP>/loader/entries/`.
A name matching `nixos-*.conf` is REJECTED by an assertion: the
installer garbage-collects loader entries with exactly that regex,
so such an entry would be deleted on the next deploy.
'';
};
};
coreutilsPackage = mkPackageOption pkgs "coreutils" { };
rsyncPackage = mkPackageOption pkgs "rsync" { };
utilLinuxPackage = mkPackageOption pkgs "util-linux" { };
};
options.boot.zfsMountBarriers = mkOption {
type = types.attrsOf barrierType;
default = { };
example = lib.literalExpression ''
{
tank = {
mountpoint = "/tank";
importUnits = [ "zfs-import-tank.service" ];
};
}
'';
description = ''
Oneshot units named `wait-for-zfs-<name>.service` that block until a ZFS
mountpoint is really mounted. Consumers order themselves with
`after`/`requires` on the barrier instead of racing an unmounted path.
'';
};
config = mkMerge [
(mkIf cfg.enable {
assertions = [
{
assertion = sdb.enable;
message = ''
boot.mirroredEsp requires boot.loader.systemd-boot.enable — the
recovery restore and the ESP mirror both run from
boot.loader.systemd-boot.extraInstallCommands.
'';
}
{
assertion = !cfg.mirror.enable || cfg.mirror.mountPoint != cfg.primary.mountPoint;
message = "boot.mirroredEsp.mirror.mountPoint must differ from boot.mirroredEsp.primary.mountPoint (rsync --delete onto itself).";
}
{
assertion =
if sdb.xbootldrMountPoint != null then
cfg.primary.mountPoint == sdb.xbootldrMountPoint
else
cfg.primary.mountPoint == config.boot.loader.efi.efiSysMountPoint;
message = ''
boot.mirroredEsp.primary.mountPoint (${cfg.primary.mountPoint}) must be
the partition systemd-boot writes entries and kernels to. With
boot.loader.systemd-boot.xbootldrMountPoint set that is the XBOOTLDR
mount, otherwise it is boot.loader.efi.efiSysMountPoint.
'';
}
{
assertion = badEntryNames == [ ];
message = ''
boot.mirroredEsp.recovery.entryFiles must not match nixos-*.conf
(offending: ${concatStringsSep ", " badEntryNames}).
systemd-boot-builder.py's garbage_collect() deletes every
loader/entries file matching the regex `nixos-.+\.conf` that is not
owned by a live generation, so such a recovery entry disappears on
the next deploy. Rename it, e.g. recovery-shell-init.conf.
'';
}
{
assertion =
!cfg.recovery.enable || !(lib.hasPrefix (cfg.primary.mountPoint + "/") cfg.recovery.directory);
message = ''
boot.mirroredEsp.recovery.directory (${cfg.recovery.directory}) is inside
the ESP. The installer garbage-collects <ESP>/EFI/nixos, so a payload
stored there is not pinned — it is exactly the thing being protected
against. Stage it on persistent non-ESP storage.
'';
}
];
warnings =
lib.optional
(cfg.recovery.enable && cfg.recovery.efiFiles == [ ] && cfg.recovery.entryFiles == [ ])
"boot.mirroredEsp.recovery.enable is on but both efiFiles and entryFiles are empty; nothing is pinned.";
boot.loader.systemd-boot.extraInstallCommands = installScript;
fileSystems = mkMerge [
(mkIf (cfg.primary.device != null) {
${cfg.primary.mountPoint}.device = mkForce cfg.primary.device;
})
(mkIf (cfg.mirror.enable && cfg.mirror.device != null) {
${cfg.mirror.mountPoint}.device = mkForce cfg.mirror.device;
})
(mkIf (cfg.mountOptions != [ ]) {
${cfg.primary.mountPoint}.options = cfg.mountOptions;
})
(mkIf (cfg.mirror.enable && cfg.mountOptions != [ ]) {
${cfg.mirror.mountPoint}.options = cfg.mountOptions;
})
];
})
{
systemd.services = lib.mapAttrs' (
name: barrier: lib.nameValuePair "wait-for-zfs-${name}" (barrierUnit barrier)
) config.boot.zfsMountBarriers;
}
];
}