Skip to content

nvidia-docker-gpu

Modules

A NixOS module that makes docker run --gpus all ... actually work with an NVIDIA GPU — as a single import, not a half-configured trap.

The problem

GPU passthrough to Docker on NixOS is a two-part switch, and enabling only one half gives you a Docker daemon that starts fine but silently can't see the GPU. --gpus all resolves to nothing, and the failure looks like a driver or image problem rather than a config gap.

The two halves are:

  1. hardware.nvidia-container-toolkit.enable — installs the Container Device Interface (CDI) spec generator, which describes the host GPU as a nvidia.com/gpu=all device.
  2. A Docker daemon with the cdi feature gate turned on — without daemon.settings.features.cdi = true, dockerd ignores the generated CDI spec entirely, so the device that half 1 produced is never wired into containers.

It is easy to flip switch 1, assume you're done, and never realize switch 2 is the load-bearing one. This module flips both, in one place.

The traps this bakes in

  • The CDI gate must be set on each daemon. Rootless Docker runs a separate dockerd with its own config; it does not inherit the rootful daemon's settings, so the features.cdi gate is applied to both.
  • Rootless containers get no DNS by default. The rootless network stack does not inherit the host's /etc/resolv.conf the way the rootful daemon does, so containers on the rootless daemon fail to resolve names until you pin a resolver. rootlessDns does that (default 1.1.1.1; override it with your own resolver).
  • Storage-driver / filesystem mismatch stops dockerd from starting. If your Docker data-root lives on a filesystem whose graphdriver has to be chosen explicitly (e.g. ZFS), a wrong or default driver makes dockerd fail with "wrong filesystem" rather than fall back. Set storageDriver (and usually dataRoot) to match.

Usage

{
  imports = [ ./nvidia-docker-gpu ];

  virtualisation.nvidiaDockerGpu.enable = true;
}

Then, after a rebuild:

$ docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi

should list the host GPU.

Options

Option Default Purpose
enable false Turn the whole thing on.
rootless true Also run a rootless daemon (safer than adding your user to the root-equivalent docker group). The CDI gate is applied here too.
rootlessDns [ "1.1.1.1" ] DNS servers for the rootless daemon's bridge — without this, rootless containers can't resolve names.
storageDriver null Force a storage driver (e.g. "zfs") when the default won't init on your data-root's filesystem.
dataRoot null Override Docker's data-root — useful on impermanent roots so images survive a reboot.

Caveats

  • This module owns virtualisation.docker. If you already configure Docker elsewhere, merge the two or drop the overlapping settings — otherwise you'll get option conflicts.
  • You still need the NVIDIA driver itself enabled (hardware.nvidia / services.xserver.videoDrivers), plus the GPU actually attached to the host (not passed through to a VM). This module handles the container plumbing, not the driver.
  • Requires a NixOS new enough to expose hardware.nvidia-container-toolkit (NixOS 24.05+, where it superseded the deprecated virtualisation.docker.enableNvidia), and a Docker new enough to honour CDI (25+).
  • The module also sets virtualisation.docker.enableOnBoot = true, so the rootful daemon starts at boot rather than on socket activation.
  • dataRoot is written into the rootful daemon's settings only; the rootless daemon keeps its own per-user data root.

Source

modules/nvidia-docker-gpu/default.nix
# nvidia-docker-gpu — GPU passthrough to Docker on NixOS
#
# GPU passthrough to `docker run --gpus all ...` on NixOS is a TWO-PART switch,
# and enabling only one half silently gives you a Docker that can't see the GPU:
#
#   1. hardware.nvidia-container-toolkit.enable — installs the CDI spec generator
#      that describes the host GPU as a Container Device Interface device.
#   2. A Docker daemon with the `cdi` feature gate turned ON — without it dockerd
#      ignores the generated CDI spec, and `--gpus` / `--device nvidia.com/gpu=all`
#      resolve to nothing.
#
# A common shape for this is a one-line "joint enable" module that flips both
# switches but delegates the load-bearing daemon config to a separate Docker
# module — and then that daemon config is easy to get wrong or forget. This
# module INLINES the daemon config (CDI gate + the rootless-DNS fix), so
# importing this single module is genuinely enough to get working GPU containers.
{
  config,
  lib,
  ...
}:
let
  cfg = config.virtualisation.nvidiaDockerGpu;
in
{
  options.virtualisation.nvidiaDockerGpu = {
    enable = lib.mkEnableOption "Docker configured for NVIDIA GPU passthrough (CDI)";

    rootless = lib.mkOption {
      type = lib.types.bool;
      default = true;
      description = ''
        Also run a rootless Docker daemon for the invoking user. Rootless is the
        safer default (no docker-group == root-equivalent handout), and the CDI
        feature gate is applied to the rootless daemon too.
      '';
    };

    rootlessDns = lib.mkOption {
      type = lib.types.listOf lib.types.str;
      default = [ "1.1.1.1" ];
      description = ''
        DNS servers for the ROOTLESS daemon's default bridge network. The trap:
        the rootless network stack does not inherit the host's /etc/resolv.conf
        the way the rootful daemon does, so containers on the rootless daemon get
        no working resolver unless you pin one here. Use your own resolver if you
        do not want to hardcode a public one.
      '';
    };

    storageDriver = lib.mkOption {
      type = lib.types.nullOr lib.types.str;
      default = null;
      example = "zfs";
      description = ''
        Force a specific Docker storage driver. Leave `null` to let NixOS pick.
        Set this (e.g. "zfs") when Docker's data-root lives on a filesystem whose
        graphdriver must be chosen explicitly — a mismatched driver makes dockerd
        fail to start ("wrong filesystem") rather than fall back gracefully.
      '';
    };

    dataRoot = lib.mkOption {
      type = lib.types.nullOr lib.types.str;
      default = null;
      example = "/persist/docker";
      description = ''
        Override Docker's data-root. Useful on impermanent / rollback roots where
        `/var/lib/docker` would be wiped on boot — point it at a durable dataset.
      '';
    };
  };

  config = lib.mkIf cfg.enable {
    # Half 1: generate the CDI spec describing the host GPU.
    hardware.nvidia-container-toolkit.enable = true;

    # Half 2: a Docker daemon that actually honours that CDI spec.
    virtualisation.docker = {
      enable = true;
      enableOnBoot = true;

      # THE load-bearing line: without the cdi feature gate, dockerd ignores the
      # CDI device spec and `--gpus all` sees no GPU.
      daemon.settings = {
        features.cdi = true;
      } // lib.optionalAttrs (cfg.dataRoot != null) {
        data-root = cfg.dataRoot;
      };

      storageDriver = lib.mkIf (cfg.storageDriver != null) cfg.storageDriver;

      rootless = lib.mkIf cfg.rootless {
        enable = true;
        setSocketVariable = true;
        daemon.settings = {
          # The gate must be repeated for the rootless daemon — it is a separate
          # dockerd with its own config, it does not inherit the rootful one.
          features.cdi = true;
          # ...and its own DNS, or rootless containers cannot resolve names.
          dns = cfg.rootlessDns;
        };
      };
    };
  };
}