· · 10 min read

Isolating a Nix remote builder in a microVM on a shared host

Distributed builds lend spare capacity to CI, but build scripts execute untrusted code. I isolate them in a builder VM with its own kernel, restricted access and resources the host can reclaim.

Send builds where the capacity is

Distributed builds let CI use another machine’s CPU, memory and disk to compile builds and run software tests. Nix sends build recipes, called derivations, and their inputs over SSH, then copies the results back. My CI controller has four cores and a small disk; this became a problem when builds filled it and stopped CI mid-compile.

Some tests drive equipment wired to the controller, so the CI runner stays there. I send hardware-independent builds, including software checks, to a second machine, the host. They run inside a VM we’ll call the builder guest. Sending only that work avoids relocating the equipment or arranging VM access to it.

The guest gives untrusted build code its own kernel. The host also has other work to do, so it lends the guest capped CPU and memory and can stop it to reclaim those resources. On NixOS, nix.distributedBuilds and a nix.buildMachines entry configure the controller’s side; the guest and its limits are NixOS configuration too.

All of that is configuration, and that’s the advantage of a declarative operating system. On NixOS I don’t set a machine up by hand and hope it stays that way; I write down the state I want, and the system is built from that description. One repository describes every machine here, from high-level fleet command, meaning which machine sends its builds to which, down to low-level resource management on a single host, such as the guest’s memory cap and the hotkey that takes the machine back. Lending a machine to CI is one versioned change to that description, not hand edits scattered across each box.

Give untrusted builds a separate kernel and a restricted account

Building a change runs that change’s own code: its build scripts, code generators and tests. Nix’s sandbox isolates builds using Linux namespaces, but shares the machine’s kernel. A kernel exploit could escape that sandbox and take over the host.

The guest is a microvm.nix VM, configured through NixOS. It runs under QEMU with its own kernel and no passed-through devices, and it reads the host’s Nix store, the directory of build inputs and results, through a read-only share. If a build becomes compromised by a kernel exploit, it now only lands in the guest. To reach the host it would need a second bug, in one of the pieces the guest touches: the host kernel’s KVM and network path, the host process that serves the shared store, or the host-side mount used when I repair the guest’s disk.

A persistent builder guest on a shared host.The CI controller runs the CI runner and keeps the attached test equipment. Build requests and inputs travel over SSH, forwarded by the host, to one persistent build VM on the host; results return to the controller. The guest has its own kernel, an untrusted build account, and persistent writable storage with its own Nix database. The host shares its whole store read-only, and the guest imports the path records at startup. The guest is not on the private network: from the private network, only the controller can reach its forwarded SSH endpoint. New guest-initiated connections are limited to web traffic and two fixed DNS resolvers, with private destinations blocked; replies to established connections remain allowed.CI controllerFour cores, small diskCI runnerAttached test equipmentHostHas its own work to doPersistent build VMOwn kernelUntrusted build accountOut: DNS and the web onlyOwn writable diskOverlay + separate Nix databaseBuild requests + inputsSSH forwarded by hostResults copied backHost's existing Nix storeThe whole store, shared in placeRead-only filesPath records imported at start
A persistent builder guest on a shared host.The CI controller runs the CI runner and keeps the attached test equipment. Build requests and inputs travel over SSH, forwarded by the host, to one persistent build VM on the host; results return to the controller. The guest has its own kernel, an untrusted build account, and persistent writable storage with its own Nix database. The host shares its whole store read-only, and the guest imports the path records at startup. The guest is not on the private network: from the private network, only the controller can reach its forwarded SSH endpoint. New guest-initiated connections are limited to web traffic and two fixed DNS resolvers, with private destinations blocked; replies to established connections remain allowed.CI controllerFour cores, small diskCI runner + attached test equipmentHostHas its own work to doPersistent build VMOwn kernelUntrusted build accountOut: DNS and the web onlyOwn writable diskOverlay + separate Nix databaseBuild requests+ inputsSSH forwarded by hostResultsbackRead-only filesPath records importedat guest startHost's existing Nix storeThe whole store, shared in place
Hardware-facing CI stays on the controller; ordinary builds try the persistent builder guest first. From the private network, only the controller can reach the guest’s forwarded, key-only SSH endpoint.

Spawning a fresh VM with a blank writable state for each build would limit persistence after a compromise. I choose to keep one persisting guest to reuse earlier build results; this drastically improves build times, however a compromised guest could tamper with later builds.

The account the controller connects as has no sudo and stays outside trusted-users, Nix’s privileged client list. It can open a shell and request builds but cannot change restricted settings, such as adding binary caches. Adding it to trusted-users would grant essentially root access to the guest. Builds run under separate sandbox accounts either way. The guest keeps require-sigs enabled for signature checks on paths the controller sends it.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
users.users."<build-account>" = {
  isNormalUser = true;
  hashedPassword = "!";                 # no password login
  openssh.authorizedKeys.keys = [ "<automation public key>" ];
  # ...
};
security.sudo.enable = false;
services.openssh.settings = {
  AllowUsers = [ "<build-account>" ];
  PermitRootLogin = "no";
  PasswordAuthentication = false;
  KbdInteractiveAuthentication = false;
  AuthenticationMethods = "publickey";
  DisableForwarding = true;
};
nix.settings = {
  allowed-users = [ "<build-account>" "<superuser>" ];
  trusted-users = [ "<superuser>" ];           # the build account stays untrusted
  sandbox = true;
  require-sigs = true;                 # default, shown explicitly
};

Guest configuration excerpt. Account names and the public key are placeholders; the signature-checking default is explicit.

One way in, DNS and the web out

On the private network, the host exposes only the guest’s forwarded, key-only SSH endpoint, reachable only from the CI controller. The host runs no SSH daemon. New outbound guest connections can reach two fixed DNS resolvers and public web addresses, for source downloads and prebuilt results. Private destinations, IPv6 and host services are blocked.

The host installs the guest’s firewall rules atomically. The (guest) VM’s systemd unit is bound to the service that installs those rules, so the guest can’t run without its firewall. The generated rules were tested in isolated network namespaces. Connection checks from the controller and inside the guest covered permitted and blocked traffic.

Put together, the guest can do exactly one job. The controller can reach it to request builds, it can fetch what builds need from the web, and nothing else on the network can reach it, nor it them. It can still send data out over the web, so this is a boundary around what builds can touch, not around what they can say.

Lend resources the host can take back

The host has a 14-core, 20-thread CPU and 16 GB of RAM. Idle cores are cheap to lend. Memory is scarce: every parallel compile job needs its own. When a kernel build had all 20 threads, it thrashed instead of compiling (more build threads mean more memory, not just more CPU); capping the guest at 14 virtual CPUs seemed to ease that. The guest gets 10 GB of RAM; its hypervisor has a 12 GB memory ceiling and 14 CPUs’ worth of execution time. That CPU quota reserves no cores, so a saturated build still competes with host work.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
systemd.services."microvm@build-guest" = {
  after = [ "build-guest-network.service" ];  # install rules before startup
  bindsTo = [ "build-guest-network.service" ]; # stop if the rule service stops
  serviceConfig = {
    CPUQuota = "1400%";      # ceiling for QEMU and its worker threads
    MemoryMax = "12G";
    MemorySwapMax = "12G";
    NoNewPrivileges = true;
    ProtectSystem = "strict";
    # ...
  };
};

The host side: the (guest) VM bound to its network rules, and the hypervisor’s limits. Names edited for publication.

Even capped, a busy guest can hold 10 of the host’s 16 GB, so I need a way to take it back. Taking the machine back costs one keypress: a hotkey on the host’s keyboard stops the guest. Interrupted builds need a retry. The host gets CPU and memory back, but not disk: the guest’s image stays (a low cost, disk is cheap).

The guest VM is started manually; simple updates/rebuilds of the host’s NixOS configuration leave a running guest alone. A running build survives a routine host rebuild; one that changes the guest’s network rules can still stop it. A nixpkgs update can do the same, because it changes the tools that service runs, which counts as changing the service.

Stopping isn’t the only way to take memory back. A setting makes the same keypress freeze the guest instead: systemctl freeze on the guest pauses its processes, and a write to the cgroup’s memory.reclaim asks Linux to push its memory into swap. I keep it off for the extra moving parts and upkeep, like a controller SSH keepalive window long enough to outlast the freeze. A freeze holds the controller’s build and the host’s swap while it lasts, so I’d only turn it on when a long uncached build needs a short pause.

Make fallback an explicit policy

Ordinary builds try the guest first. If it is stopped or every build slot is busy, they may fall back to the controller’s own kernel, outside the VM’s protection, because the host isn’t dedicated and CI shouldn’t stall whenever it’s needed for other work. That is the exposure the controller had before the guest existed: the guest protects the host, not the controller. Builds that Nix prefers to run locally never leave the controller at all.

One build is the exception and must not fall back. It compiles against a kernel whose build is cached only in the guest; on the controller, Nix would first rebuild that entire kernel on four cores and a small disk. So that build doesn’t go through Nix’s routing. A small script sends it to the guest and refuses to run it anywhere else, including when the guest or its cached kernel is unavailable. Failing loudly there beats a silent kernel rebuild on the wrong machine.

The guest isolates the build; it doesn’t vouch for the result. The controller trusts whatever comes back, including test verdicts.

The rule I’d carry to any setup: keep the CI work that needs hardware next to the hardware, send builds where cores and disk are cheap, cap their memory, and put them behind a separate kernel if that machine matters.

Share