Why per-caller locks didn’t compose
Shared test hardware needs exclusive access, and the usual tool is a lease, a time-limited grant of the device to one holder. It comes paired with a fencing token, a number that rises with every grant, so the device can refuse commands from a holder whose turn is over. I haven’t evaluated labgrid or LAVA, which share boards across whole labs; what I needed was narrower: a way to take one board’s login away from every caller but one service.
My case is one small test board built around the SG2000 SoC, which CI jobs, AI coding agents and I were all using to test display-driver changes. I can’t replace it. The rest of the setup is one USB capture device that records its HDMI output and an old laptop wired to both, which I’ll call the controller. Nothing in the automation can power-cycle the board, so a hung board waits for me. Each kind of caller already had a lock of its own.
A lock file in /tmp kept one test script’s runs apart, but a service running with systemd’s PrivateTmp never saw it. Marker files in the board’s /run covered only commands on the board, and only for one boot. A CI concurrency group kept CI jobs off each other. The agents ran as coding sessions launched from issues in my tracker, which I’ll call workers. A cap of one worker kept them apart. Meanwhile the callers’ tools carried the board’s password themselves, some as a built-in default. Three accounts could also open the capture device directly: mine, the agents’ service account and the CI runner’s. Every lock serialized its own kind of caller and nobody else.
My requirement was about attention. A wedged board could call me in, but the common failures shouldn’t need me at all. The obvious patch is one shared lock for every caller. A cooperative lock can’t promise exclusivity, though, while another caller holds the board’s login to walk around it.
Give the board’s login to one service
The broker is a service on the controller, and the test tools were rewritten to send their board commands through it instead of carrying the password. In the controller’s configuration, the board’s password and the broker’s SSH key for it sit in root-only secret files. Systemd passes a copy only to the broker’s service, as systemd credentials, so in that configuration the broker is the only automated caller that holds the board’s login. It keeps its state and a ledger of every operation in SQLite, written before each step runs, so a restart forgets nothing. Its clients fail closed: no broker, no board.
That login boundary has gaps I haven’t closed. The board image’s own configuration, in the repository the workers check out, declares a default password for the board’s user. The hardware CI runner can still open the board’s serial console. And the capture device belongs to the broker’s group, which that runner and the agents’ service account both join to reach the broker’s socket, so either could open the device directly.
AI sessions wrote, reviewed and committed nearly all of the broker; its design came from plans they drafted at my request, which I told them to build. The broker is a state machine with six states: available, reserved, executing, verifying (checking a finished change), quarantined and recovering. Only an available board takes a new holder. A reserved lease hasn’t changed anything on the board yet, so the broker now reclaims it from a caller that died. There is no timer for that, because a timeout can’t prove a command on the board stopped. The broker records the caller’s cgroup when it grants the lease and reclaims only once that cgroup is empty or gone. An executing lease has a change in flight, so the broker never reclaims it after a crash. The board goes into quarantine instead: nothing may use it until someone knows what state it’s in, and no timeout clears that.
A lease with two fencing tokens, taken only to touch the board
The lease covers every step that touches the board, from copying files onto it to the last capture, read-only checks included. Writing the patch and building the modules happen off the board and need no lease, so a worker can prepare its next attempt while another holds the board or while the board sits in quarantine.
A command formed under an earlier lease or before a reboot can still arrive late, so each lease carries two fencing tokens. The first is a generation counter the broker bumps on every acquire, every release and every confirmed reboot; a command carrying an old generation is refused. The second is the board’s own boot id, the random identifier Linux picks at each boot. It doesn’t rise like the counter, but it changes at every boot, which is all this fence needs. It also catches a reboot the broker never confirmed. A command can name the boot it was formed for, and every reboot and every command from the trial runner does. The board re-reads its boot id just before such a command runs, so a command formed against one boot can’t land on the next. A command sent without one skips that check.
The usual objection to fencing tokens is that they only work if the resource checks them. The boot id is re-read on the board itself, but only inside the command the broker sends. The broker checks the generation. Both work only while no other automated caller holds the login. That’s why the gaps above matter: a fence binds only the callers that pass through it, and one that reaches the board some other way skips both checks.
Quarantine only when the outcome is unknown
Fencing can’t say what a command that lost its connection did to the board. That is the case quarantine exists for, and getting it to behave took three rounds.
First, a caller’s word could clear it. A caller releases its lease either clean or in doubt. Before the broker touched the board, an AI session’s adversarial review of the code found that a holder’s own clean release cleared a quarantine and freed the board. The main run script always released clean. That script’s default path also never looked up the board’s address, so every run would have quarantined the board without reaching it, and the same clean release would have hidden that. Both were fixed.
Second, the design still quarantined on outcomes that were safe: a script on the board that refused and changed nothing, a command that ran to the end and failed, a capture that timed out, and a reserved lease the broker found after a restart. None meant the board was mid-change. The restart case quarantined because the broker couldn’t confirm the lease’s owner was alive, and a configuration switch can trigger it. I hit it myself, on a read-only check. I objected that if every bad run meant I had to clear a quarantine, the design put me back at the bench. I said the broker could clear one itself once a reboot worked, and asked for the plan to be reviewed for other steps that would call me back.
I told a session to build the revised plan, and the next commit to the broker narrowed the rule. A remote command that runs to completion is a known outcome, whatever it returns. Only a lost connection with no output, or a command that timed out, is uncertain. A refusal returns the broker to the state it was in before the operation. A capture timeout no longer counts as uncertain, because capturing doesn’t change the board. A restart reclaims a lease that changed nothing. Release is judged from the broker’s own ledger of operations, not from what the caller says. A caller can still force a quarantine when it releases, but a clean release can never clear one.
Third, a caller could still force a quarantine, and the first driver trial through the broker ended that way. Its one reboot came back. Then a check on the fresh boot refused, harmlessly, because the marker the loader leaves once the new driver loads was missing, most likely from a loader race that a later offline test reproduced. The script running the trial counted the unconfirmed trial as hardware doubt and forced a quarantine when it released. I authorized a clear, and after a loader fix and a capture fix, a manual run through the broker released the board clean. A caller that can force a quarantine carries its own copy of the rule, so narrowing the rule means narrowing it there too.
Send at most one guarded reboot per boot
Rebooting is the obvious way out of an unknown state, but a reboot can leave the board unable to answer at all. The reboot rules come from an outage before the broker existed. A CI job soft-rebooted the board while it ran an older NixOS image on the vendor’s kernel. That image had no reliable restart path. Only systemd’s runtime watchdog was armed, not its reboot watchdog, so a reboot that stalled had nothing to reset the SoC. The board never came back, and it sat unreachable for two days until I power-cycled it. CI kept running hardware jobs against the dead board, and main went red.
The first fix only contained the damage: a switch file committed to the board’s repository made those hardware steps skip with a warning, so a hung board no longer turns main red. The switch is on again because the hardware workflows target that older image, while the board now runs a mainline kernel. Pull-request CI checks software only. Hardware evidence comes from runs that hold a lease, and those leases have gone to workers and to me.
I still wanted a soft reboot wherever the workflow needed one, rather than unplugging cables for every restart. The broker allows them only under two rules. First, a table of the board’s installed images records each one’s SSH host key, which the broker checks to confirm which image it reached. The table also records whether an image’s reboots reset the SoC through the hardware watchdog, as the mainline image’s do without relying on systemd’s reboot watchdog. The broker refuses a soft reboot on any image the table doesn’t mark that way, or doesn’t list. Second, it sends at most one guarded reboot per boot.
A guard on the board checks the expected boot id, kernel, the exact NixOS system it booted, an idle display pipeline and the watchdog restart provider, then leaves a once-per-boot marker and schedules the reboot. The broker then waits up to ten minutes for the board to pass the same checks on a new boot id, which I’ll call a verified idle boot. If none appears, it never sends a second reboot: the outcome counts as uncertain, the broker records it and quarantines the board, and any physical recovery is mine. Two recorded boots took about four and a half minutes to reach SSH, and an earlier four-minute window once ran out while the board was still booting. A wait that ends early makes a slow boot look like a hung one.
Add workers once the broker can refuse a second holder
Because many tasks never touch the board, I asked for a second worker, but only once the broker could refuse a second holder on its own. Software-only work could then run beside a live measurement, and a second hardware task would wait or be refused.
No worker is allowed to clear a quarantine. That rule lives in the instructions each worker starts with, not in a check in the broker. So it is the kind of cooperative rule the broker replaced, and one more gap. A clear takes my say-so, which the broker records, or the broker’s own recovery, which may reboot the board once and clears the quarantine only on a verified idle boot. If it can’t get one, it files an issue in my tracker asking me for a hand. That path is built, held off by a flag, and has never run live. Clearing only on that boot means recovery accepts no state the normal loop wouldn’t.
The broker is the only service on the controller given the board’s login, and the gaps above are still open. When several kinds of caller share one scarce device, put its credential in one service’s configuration and nowhere else, and close every other way in. A lease, fencing tokens and more workers only mean something after that.