etcd-cluster-over-tailnet¶
Modules
A NixOS module that runs an etcd cluster whose peer and client traffic rides only a private mesh interface (Tailscale, WireGuard, or a private VLAN) and is never exposed on the public firewall. All etcd URLs are derived from a single topology attrset, and the voting set is kept deliberately narrow and co-located so a WAN link flap can't cost you quorum.
The problem¶
etcd is a Raft-based distributed store — often the DCS backing something like
Patroni's PostgreSQL leader election. Two things make a naive services.etcd
setup fragile:
-
Quorum is latency- and partition-sensitive. If you spread etcd voters across sites connected by a flaky WAN link, a single transatlantic hiccup can drop the cluster below quorum or trigger a spurious leader election — exactly when you least want it.
-
etcd URLs are easy to get wrong. Client/peer/advertise URLs, the initial-cluster string, ports, and the cluster token all have to agree across every member. Duplicating those literals per host rots quickly.
The key insight / trap¶
-
Keep the voting set small and in one low-latency zone. Machines can consume etcd (or run Patroni/replicas) without being etcd voters. Only list the co-located members in
peers. A cross-WAN node that never votes can't break quorum when its link flaps. This module makes the voting set an explicit option so the boundary is obvious and one-line to change. -
One topology, many derived URLs. The
peersattrset (name → private-mesh address) is the entire topology.listenClientUrls,listenPeerUrls,advertiseClientUrls,initialAdvertisePeerUrls, andinitialClusterare all computed from it. Adding or removing a voter is a one-line edit, and every member computes the identicalinitialClusterlist. -
Mesh-only, per-interface firewall. The client and peer ports are opened with
networking.firewall.interfaces.<iface>.allowedTCPPorts, i.e. only on the mesh interface. They stay closed on every public NIC regardless of your global firewall policy. Peer traffic is never advertised on a public address. -
A loopback client URL is added on purpose so local
etcdctlworks without routing over the mesh. -
dataDirmust survive reboots. On impermanence / wiped-root hosts, pointdataDirat a persistent path (e.g./persist/etcd) and persist it — losing the raft log forces a re-bootstrap and can break the cluster. -
initialClusterStateis the one operational knob (see below).
Usage¶
Import the module and configure it identically on every voter — it only differs
by which host it runs on, and nodeName defaults to the hostname:
{
imports = [ ./etcd-cluster-over-tailnet ];
services.etcdMesh = {
enable = true;
interface = "tailscale0"; # your private mesh interface
clusterToken = "my-etcd-cluster"; # same on every member
dataDir = "/persist/etcd"; # persistent if you wipe root
peers = { # the voting set, mesh addresses
node-a = "100.100.0.1";
node-b = "100.100.0.2";
node-c = "100.100.0.3";
};
};
}
Options¶
| Option | Default | Purpose |
|---|---|---|
enable |
false |
Turn the module on. |
nodeName |
config.networking.hostName |
This member's etcd name; must be a key of peers unless nodeAddress is set. |
peers |
(required) | Voting set: member name → address on the private mesh interface. |
nodeAddress |
null |
Explicit mesh address for a node that is not in the bootstrap peers set (a member being added with initialClusterState = "existing"). null looks it up from peers. |
interface |
"tailscale0" |
Mesh interface; ports are opened only here. |
clientPort |
2379 |
etcd client API port. |
peerPort |
2380 |
etcd peer (raft) port. |
clusterToken |
"etcd-cluster" |
Shared initial-cluster-token; identical on all members. |
dataDir |
"/var/lib/etcd" |
Raft log + data dir; make persistent under impermanence. |
initialClusterState |
"new" |
"new" for bring-up, "existing" to join a live cluster. |
Bootstrapping vs. adding a member¶
-
Initial bring-up: leave
initialClusterState = "new"on all voters and bring them up together. -
Adding a voter to a running cluster: flip the new node to
initialClusterState = "existing", but only after registering it on the existing members first:
Starting a fresh node with "new" against a live cluster, or with
"existing" before the member add, makes etcd refuse to join.
A joining node is deliberately not in the static bootstrap peers set, so
it has no address to look up — give it nodeAddress explicitly (an assertion
catches the case where neither peers nor nodeAddress supplies one).
Caveats¶
- Traffic between members is plain HTTP — it relies on the mesh interface for confidentiality and authenticity. Don't route it over anything but your encrypted mesh (that's the whole point). If you need TLS between members, extend the module to set etcd's peer/client cert options.
- No client auth: every mesh peer is a full read/write admin. etcd runs with
no TLS client certs and no RBAC, so the security boundary is exactly "who can
reach the client/peer ports on the mesh". Any host on the interface can run
etcdctlagainst the client port with no credentials and rewrite any key — including keys that back things like Patroni leader election. On a flat Tailscale tailnet every node is authorized equally, so you MUST restrict reach with node ACLs (limit which peers can hit these ports). For any multi-tenant or low-trust mesh, don't rely on the network alone: enable etcd RBAC and peer/client TLS (extend the module to set etcd's cert/auth flags). clusterTokennamespaces the cluster; it is not a cryptographic secret, but keep it distinct per cluster so a stray peer can't join the wrong one.- An odd number of voters (3 or 5) is strongly recommended for quorum math.
Source¶
modules/etcd-cluster-over-tailnet/default.nix
# etcd-cluster-over-tailnet
#
# Run an etcd cluster whose peer + client traffic rides ONLY a private mesh
# interface (WireGuard / Tailscale / a VLAN) and never touches the public
# firewall. Every etcd URL is derived from a single topology attrset so that
# adding or removing a voter is a one-line edit, and the voting set is kept
# narrow and co-located so a WAN link flap can't cost you quorum.
#
# This is a generalised, self-contained NixOS module. Drop it into your config
# and import it. No external topology file is required — the topology lives in
# the `peers` option below.
#
# Usage (identical module on every voter, differing only in which host it runs
# on — `nodeName` defaults to the hostname):
#
# imports = [ ./etcd-cluster-over-tailnet ];
#
# services.etcdMesh = {
# enable = true;
# interface = "tailscale0"; # your private mesh interface
# clusterToken = "my-etcd-cluster"; # shared by all members
# peers = { # the *voting* set only
# node-a = "100.100.0.1";
# node-b = "100.100.0.2";
# node-c = "100.100.0.3";
# };
# };
#
# The `peers` attrset is the whole cluster topology: keys are etcd member
# names, values are the address each member is reachable at *on the private
# interface*. Keep this set small (3 or 5) and inside one low-latency zone.
# Machines that consume the cluster but must never vote (e.g. cross-WAN
# read replicas) simply are NOT listed here.
{
config,
lib,
...
}:
let
cfg = config.services.etcdMesh;
# This node's private-interface address. Normally looked up from the topology
# (`peers`), but a node that is JOINING an already-running cluster with
# `initialClusterState = "existing"` is registered out-of-band via
# `etcdctl member add` and is deliberately NOT part of the static bootstrap
# `peers` set — such a node advertises its address via `nodeAddress`.
nodeAddress = if cfg.nodeAddress != null then cfg.nodeAddress else cfg.peers.${cfg.nodeName};
# initial-cluster string, e.g. "node-a=http://10.0.0.1:2380,node-b=..."
# Derived from the single `peers` attrset so topology lives in exactly one
# place. Every member computes the *same* list.
initialCluster = lib.mapAttrsToList (
name: addr: "${name}=http://${addr}:${toString cfg.peerPort}"
) cfg.peers;
in
{
options.services.etcdMesh = {
enable = lib.mkEnableOption "etcd cluster node bound to a private mesh interface";
nodeName = lib.mkOption {
type = lib.types.str;
default = config.networking.hostName;
description = ''
This member's etcd name. Must be a key of `peers`. Defaults to the
machine's hostname.
'';
};
peers = lib.mkOption {
type = lib.types.attrsOf lib.types.str;
example = {
node-a = "100.100.0.1";
node-b = "100.100.0.2";
node-c = "100.100.0.3";
};
description = ''
The etcd voting set: member name -> address reachable on the private
mesh interface. This IS the cluster topology; keep it small (3/5) and
co-located in one low-latency zone so a WAN partition to any other
machine cannot break quorum. Nodes that consume etcd but must never
vote are deliberately omitted.
'';
};
nodeAddress = lib.mkOption {
type = lib.types.nullOr lib.types.str;
default = null;
example = "100.100.0.9";
description = ''
This node's address on the private mesh interface. Leave null to look it
up from `peers` (the normal case for a bootstrap voter). Set it
explicitly only for a node that runs etcd and advertises itself but is
NOT in the static bootstrap `peers` set — i.e. a member being added to an
already-running cluster with `initialClusterState = "existing"` after an
out-of-band `etcdctl member add`.
'';
};
interface = lib.mkOption {
type = lib.types.str;
default = "tailscale0";
example = "wg0";
description = ''
Private mesh interface name. etcd's client and peer ports are opened
ONLY on this interface — never on the public firewall.
'';
};
clientPort = lib.mkOption {
type = lib.types.port;
default = 2379;
description = "etcd client API port.";
};
peerPort = lib.mkOption {
type = lib.types.port;
default = 2380;
description = "etcd peer (raft) port.";
};
clusterToken = lib.mkOption {
type = lib.types.str;
default = "etcd-cluster";
description = ''
Shared initial-cluster-token. Every member must use the same value;
it namespaces the cluster so a stray peer from another cluster can't
accidentally join. Not a secret in the cryptographic sense, but keep
it distinct per cluster.
'';
};
dataDir = lib.mkOption {
type = lib.types.str;
default = "/var/lib/etcd";
example = "/persist/etcd";
description = ''
Raft log + data directory. If you run an impermanence / wiped-root
setup, point this at a persistent path (and persist it) so the raft
log survives a reboot — losing it makes the node re-bootstrap and can
break the cluster.
'';
};
initialClusterState = lib.mkOption {
type = lib.types.enum [
"new"
"existing"
];
default = "new";
description = ''
etcd initial-cluster-state — the one operational knob.
Leave "new" for the initial bring-up of all voters at once.
Flip to "existing" when adding a voter to an already-running cluster,
and ONLY after you have registered the new peer on the existing members
with `etcdctl member add <name> --peer-urls=http://<addr>:<peerPort>`.
Starting a fresh node with "new" against a live cluster, or with
"existing" before the member-add, makes etcd refuse to join.
'';
};
};
config = lib.mkIf cfg.enable {
assertions = [
{
assertion = (cfg.peers ? ${cfg.nodeName}) || (cfg.nodeAddress != null);
message = "services.etcdMesh: nodeName \"${cfg.nodeName}\" is not a key of `peers` and no explicit `nodeAddress` was set.";
}
];
services.etcd = {
enable = true;
name = cfg.nodeName;
dataDir = cfg.dataDir;
# Bind the client API to the mesh address AND loopback, so local
# `etcdctl` works without going over the network.
listenClientUrls = [
"http://${nodeAddress}:${toString cfg.clientPort}"
"http://127.0.0.1:${toString cfg.clientPort}"
];
# Peer (raft) traffic is mesh-only — never advertise a public address.
listenPeerUrls = [
"http://${nodeAddress}:${toString cfg.peerPort}"
];
advertiseClientUrls = [
"http://${nodeAddress}:${toString cfg.clientPort}"
];
initialAdvertisePeerUrls = [
"http://${nodeAddress}:${toString cfg.peerPort}"
];
inherit initialCluster;
initialClusterToken = cfg.clusterToken;
initialClusterState = cfg.initialClusterState;
};
# Open the etcd ports ONLY on the private mesh interface. Because these
# are per-interface rules, the ports stay closed on every public NIC even
# if the global firewall is otherwise permissive.
networking.firewall.interfaces.${cfg.interface}.allowedTCPPorts = [
cfg.clientPort
cfg.peerPort
];
};
}