· · 11 min read

Sharing a host's Nix store with a builder VM

A build VM can read its host's Nix store through a read-only share. The host's files stay immutable; the VM's view of them doesn't. Its garbage collector can hide shared files behind overlay whiteouts unless the VM is told, on every start, that those files are still needed.

Reuse the host’s build inputs and results

My CI sends its builds to another machine over SSH, a Nix setup called distributed builds. The machine that requests them is the CI controller. The one that runs them is a VM, the builder guest. It lives on a machine with capacity to spare, the host. Why the builder is a VM, and how it’s locked down, is its own post; this one is about the store it shares.

Nix keeps sources, tools and build results in its store; each stored object has a unique location called a store path. The host’s own store may already hold what a guest build needs, such as a kernel the guest would otherwise compile again. The guest can copy those objects, or read them in place through a read-only share. Sharing avoids duplicate storage and transfers, but every shared path is then tracked by two databases and cleaned up by two garbage collectors, the host’s and the guest’s.

microvm.nix configures VMs through NixOS. Its virtiofs file share exposes the host’s store. The writableStoreOverlay option combines that share with guest storage in an OverlayFS mount: host files form the read-only lower layer, and guest writes go into the upper layer on the guest’s disk. The host’s file-sharing process enforces the read-only side, so the host’s files are immutable to the guest; the overlay on top is where the trap in the next section lives. Builds in the guest run untrusted code, and I share the whole store, so the guest can read everything there and send it over the web. No secrets belong in it.

Files alone are not enough. Nix only trusts a store path its database has a record of, so a shared file the guest’s database has never heard of is invisible to it. At each guest start, a host service exports the host’s path records into a read-only snapshot. Before the guest accepts SSH or build requests, it imports the snapshot into its own database using nix-store --load-db; the host’s live database stays private. New host outputs reach the guest’s database at its next start.

Guest cleanup can hide read-only host files

The guest’s garbage collector, Nix’s cleanup of unused paths, once ran against the imported host paths. Nothing in the guest had marked them as needed, so it “deleted” files it had no way to delete. On an overlay, deleting a lower-layer file writes a whiteout, a marker in the upper layer that hides it. The bytes stayed on the host, but the next startup import failed because a build-recipe file was hidden. SSH depended on that import, so the VM ran without accepting SSH connections.

A garbage-collection root marks a path and its dependencies as needed. My fix tells the guest that every shared file is still needed. At every start, before SSH, build requests or collection can run, it roots each path in the snapshot, then imports the records. The unit, from the guest’s config:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
  # Files alone do not make cached host outputs valid in the guest database.
  # Protect the shared lower store from guest GC as well as host GC. Otherwise
  # GC creates persistent whiteouts and the next boot's full import fails.
  # Both scheduled GC and daemon-triggered GC must wait for these guest roots.
  systemd.services.import-host-store = {
    requiredBy = [ "nix-daemon.service" "sshd.service" "nix-gc.service" ];
    before = [ "nix-daemon.service" "sshd.service" "nix-gc.service" ];
    after = [ "local-fs.target" ];
    unitConfig.RequiresMountsFor = [ "/run/host-store-metadata" "/nix/.ro-store" "/nix/store" ];
    serviceConfig = { Type = "oneshot"; RemainAfterExit = true; };
    script = ''
      ${pkgs.python3}/bin/python3 ${./pin-host-store.py} /run/host-store-metadata/registration
      ${pkgs.nix}/bin/nix-store --store local --load-db < /run/host-store-metadata/registration
    '';
  };

The pin script, which creates the roots, validates the snapshot’s record format and checks every registered path exists first. A regression test in a disposable local store, without an overlay, checks that roots protect paths through collection and reimport; it tests the roots, not the whiteouts. Removing the roots and collecting again makes the import fail. Existing whiteouts needed a reversible repair with the guest stopped; a separate script quarantined only the markers hiding paths the host still had.

I kept the writable store and its database across restarts so the guest could reuse earlier results. Resetting both at each start would discard cached results absent from the host’s shared store and remove old whiteouts, but each boot would still need roots for imported paths. That alternative remains untested here. Sharing doesn’t require a long-lived guest, though: a fresh VM per build could use the same share.

The host’s own collector could delete shared paths too, even in the moment between listing a path for the guest and protecting it. So the host roots every exported path while holding Nix’s garbage-collection lock, which keeps the collector out until the roots exist. Those roots accumulate until I wipe the guest’s disk, since retained guest results can depend on earlier imports. Until then, the host’s collector cannot free anything its store held at any guest start. The price is disk space on the host; the roots change nothing it runs.

Host roots keep the bytes; guest roots keep them visible. That holds beyond Nix: with a writable overlay, cleanup in a VM or container can hide files in a shared read-only cache.

Sign results that return as inputs

Sharing solves reuse in one direction, host to guest. Results traveling the other way raise a second problem. A result coming back from the guest is accepted on sight: Nix copies it to the controller without checking signatures. The same file going back into the guest has to prove itself. That happens when a later build needs a result the guest no longer holds. The account the controller connects as stays outside trusted-users, Nix’s privileged client list, so require-sigs applies. Most store paths are named after how they were built, not what they contain, so Nix can’t verify them by hashing; they need a signature from a key the guest trusts. Content-addressed paths carry their own hash and need none.

Here, the guest’s disk ran low and a result it had built went missing. The guest’s collector runs when a build finds the disk nearly full, so it likely removed the copy; a manual cleanup around then was never ruled out. The result lay outside the shared host store the roots protect, so this was separate from the whiteout trap. The controller pushed its copy, and the guest refused:

1
cannot add path … because it lacks a signature by a trusted key

A setup comment in the guest’s config suggested the trusted key was still a placeholder, but the running guest held the right one. The path was the problem: the controller’s copy showed "signatures": [] in nix path-info --json. Its secret-key-files setting signs locally built paths; this one was built remotely and arrived unsigned. The output hadn’t changed; its role had.

A build result returns unsigned, then is refused as a later input.The CI controller requests a build from the Build VM. The VM returns the result, which the controller accepts without a signature as the result of its requested build. Later, the VM no longer has its copy; the cause of that absence was not confirmed. The controller still has an unsigned copy and offers that same result as an input to a later build. The VM rejects the import because this input requires a signature by a trusted key. The configured response is periodic signing of the whole store on the CI controller, where the private key stays. Newly returned results can remain unsigned before a successful signing pass. Signing permits reuse under this policy; it does not check whether the build was correct. The figure describes the configured response, not a verified successful retry.CI controllerrequests buildsBuild VMbuilds softwareRequest a build1Return resultAccepted unsigned as the requested result2Controller still hasits unsigned copy3VM's copy later absentCause not confirmedSend the same result as a later input4Rejected: trusted signature required5Configured response: periodic whole-store signingOn the CI controller; the private key stays there.New results can wait unsigned until the next signing pass.Signing permits reuse; it does not check whether the build was correct.
A build result returns unsigned, then is refused as a later input.The CI controller requests a build from the Build VM. The VM returns the result, which the controller accepts without a signature as the result of its requested build. Later, the VM no longer has its copy; the cause of that absence was not confirmed. The controller still has an unsigned copy and offers that same result as an input to a later build. The VM rejects the import because this input requires a signature by a trusted key. The configured response is periodic signing of the whole store on the CI controller, where the private key stays. Newly returned results can remain unsigned before a successful signing pass. Signing permits reuse under this policy; it does not check whether the build was correct. The figure describes the configured response, not a verified successful retry.CI controllerrequests buildsBuild VMbuilds softwareRequest a build1Return resultAccepted unsigned2Controller keepsits unsigned copy3VM's copy later absentCause not confirmedSend the same resultas a later input4Rejected: trusted signature required5Configured responsePeriodic whole-store signingon the CI controller. Key stays there.New results can wait unsigneduntil the next signing pass.Signing permits reuse; it does notcheck whether the build was correct.
The same result returns unsigned, then needs a trusted signature when sent back as an input. Step 5 is the configured response; it does not resume the failed run.

A rebuild would clear this failure and a bigger disk would delay the next one; neither signs the copies the controller already holds. The fix I deployed: the guest trusts the controller’s signing key, and an hourly nix store sign --all pass on the controller covers received results too, with the private key staying there. The timer’s cost is a window: a fresh result can wait up to an hour before it can return to the guest. A hook that signs each result as it arrives would close that window for new results. It wouldn’t have signed the copies already in the store; those needed one sign --all pass regardless, and the timer is that pass made recurring.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
systemd.services.sign-store-for-offload = {
  description = "Sign the local Nix store for the builder guest";
  serviceConfig = {
    Type = "oneshot";
    ExecStart = "${pkgs.nix}/bin/nix store sign --all --key-file <signing-key-file>";
  };
};
systemd.timers.sign-store-for-offload = {
  wantedBy = [ "timers.target" ];
  timerConfig = { OnBootSec = "10min"; OnUnitActiveSec = "1h"; Persistent = true; };
};

The controller’s hourly signing pass. The key-file path is a placeholder.

Signing the whole controller store broadens what the guest accepts; it does not verify correctness. Anyone who takes over the controller could plant paths in the guest for later builds, but that machine already picks every build and trusts every result.

What Nix’s own overlay store would and wouldn’t fix

Nix has a store type built for exactly this layering, the local-overlay store. It reads which shared files are valid straight from the host’s own records, and its collector never writes a whiteout over a host file, so it would avoid the trap above. I didn’t evaluate it when I built this. It remains experimental. Layered on the live host store, it would read the host’s database or talk to its daemon from inside the VM, which is the one thing this design refuses to expose to the untrusted guest. A read-only copy of the host’s records might serve instead; I haven’t tried that. The manual also warns that Nix’s read-only database mode can return incorrect results if another process changes the database.

Both approaches retain a filesystem limit: OverlayFS requires the lower layer to remain unchanged while mounted. Host roots keep imported paths from being deleted, but the lower layer can still change whenever the host builds or collects other paths, and host rebuilds are set not to restart the guest. The kernel doesn’t promise what happens when the bottom layer changes underneath a mounted overlay, and I haven’t found out.

Share