· · 14 min read

A lease broker for one test board shared by CI and coding agents

Per-caller locks don't compose when CI jobs, coding agents and I share one fragile board. What worked better was one broker that holds the board's login, leases the board with two fencing tokens and quarantines it only when an outcome is unknown.

Why per-caller locks didn’t compose

Shared test hardware needs exclusive access, and the usual tool is a lease, a time-limited grant of the device to one holder. It comes paired with a fencing token, a number that rises with every grant, so the device can refuse commands from a holder whose turn is over. I haven’t evaluated labgrid or LAVA, which share boards across whole labs; what I needed was narrower: a way to take one board’s login away from every caller but one service.

My case is one small test board built around the SG2000 SoC, which CI jobs, AI coding agents and I were all using to test display-driver changes. I can’t replace it. The rest of the setup is one USB capture device that records its HDMI output and an old laptop wired to both, which I’ll call the controller. Nothing in the automation can power-cycle the board, so a hung board waits for me. Each kind of caller already had a lock of its own.

A lock file in /tmp kept one test script’s runs apart, but a service running with systemd’s PrivateTmp never saw it. Marker files in the board’s /run covered only commands on the board, and only for one boot. A CI concurrency group kept CI jobs off each other. The agents ran as coding sessions launched from issues in my tracker, which I’ll call workers. A cap of one worker kept them apart. Meanwhile the callers’ tools carried the board’s password themselves, some as a built-in default. Three accounts could also open the capture device directly: mine, the agents’ service account and the CI runner’s. Every lock serialized its own kind of caller and nobody else.

My requirement was about attention. A wedged board could call me in, but the common failures shouldn’t need me at all. The obvious patch is one shared lock for every caller. A cooperative lock can’t promise exclusivity, though, while another caller holds the board’s login to walk around it.

Give the board’s login to one service

The broker is a service on the controller, and the test tools were rewritten to send their board commands through it instead of carrying the password. In the controller’s configuration, the board’s password and the broker’s SSH key for it sit in root-only secret files. Systemd passes a copy only to the broker’s service, as systemd credentials, so in that configuration the broker is the only automated caller that holds the board’s login. It keeps its state and a ledger of every operation in SQLite, written before each step runs, so a restart forgets nothing. Its clients fail closed: no broker, no board.

That login boundary has gaps I haven’t closed. The board image’s own configuration, in the repository the workers check out, declares a default password for the board’s user. The hardware CI runner can still open the board’s serial console. And the capture device belongs to the broker’s group, which that runner and the agents’ service account both join to reach the broker’s socket, so either could open the device directly.

AI sessions wrote, reviewed and committed nearly all of the broker; its design came from plans they drafted at my request, which I told them to build. The broker is a state machine with six states: available, reserved, executing, verifying (checking a finished change), quarantined and recovering. Only an available board takes a new holder. A reserved lease hasn’t changed anything on the board yet, so the broker now reclaims it from a caller that died. There is no timer for that, because a timeout can’t prove a command on the board stopped. The broker records the caller’s cgroup when it grants the lease and reclaims only once that cgroup is empty or gone. An executing lease has a change in flight, so the broker never reclaims it after a crash. The board goes into quarantine instead: nothing may use it until someone knows what state it’s in, and no timeout clears that.

How the broker grants, fences and quarantines the one board.Six broker states. Focal path: available, acquire, reserved, run or reboot, executing, done, verifying, clean release, back to available. A lease spans reserved, executing and verifying and carries a generation fencing token: a stale generation or a command formed for an old boot id is refused. The deploy lock fences executing: a broker deploy waits while an operation is in flight. An uncertain outcome from executing, or an uncertain operation in the verifying ledger, moves the broker to quarantined. Automatic recovery, when the recovery-inhibit flag is absent, moves quarantined to recovering; it is built and has not run live. Recovering returns to available only on a verified idle boot; otherwise it falls back to quarantined and files a human_review recovery issue asking the operator for a hand; no live recovery or recovery ask has been recorded. Only a recorded human clear_quarantine moves quarantined straight to available; a hung board needs a power-cycle first. Not drawn: release from reserved, deliberate refusals that return to the prior state, and further boot-fenced run or reboot from verifying.lease · generation nstale token refusedacquirerun, rebootdoneclean releaseavailablereservedexecutingverifyingbroker deploydeploy lockuncertain outcomequarantinedrecovery-inhibitrecoveringauto recovernot run livefailsverified idle boothuman_reviewrecovery issuenot run liveoperatorclear_quarantinepower-cycle if hung
Each lease’s generation number fences stale commands, a broker deploy waits out any operation in flight, and an unknown outcome quarantines the board. The board leaves quarantine on a clear I authorize or through automatic recovery, held off by a flag and never run live.

A lease with two fencing tokens, taken only to touch the board

The lease covers every step that touches the board, from copying files onto it to the last capture, read-only checks included. Writing the patch and building the modules happen off the board and need no lease, so a worker can prepare its next attempt while another holds the board or while the board sits in quarantine.

A command formed under an earlier lease or before a reboot can still arrive late, so each lease carries two fencing tokens. The first is a generation counter the broker bumps on every acquire, every release and every confirmed reboot; a command carrying an old generation is refused. The second is the board’s own boot id, the random identifier Linux picks at each boot. It doesn’t rise like the counter, but it changes at every boot, which is all this fence needs. It also catches a reboot the broker never confirmed. A command can name the boot it was formed for, and every reboot and every command from the trial runner does. The board re-reads its boot id just before such a command runs, so a command formed against one boot can’t land on the next. A command sent without one skips that check.

The usual objection to fencing tokens is that they only work if the resource checks them. The boot id is re-read on the board itself, but only inside the command the broker sends. The broker checks the generation. Both work only while no other automated caller holds the login. That’s why the gaps above matter: a fence binds only the callers that pass through it, and one that reaches the board some other way skips both checks.

Quarantine only when the outcome is unknown

Fencing can’t say what a command that lost its connection did to the board. That is the case quarantine exists for, and getting it to behave took three rounds.

First, a caller’s word could clear it. A caller releases its lease either clean or in doubt. Before the broker touched the board, an AI session’s adversarial review of the code found that a holder’s own clean release cleared a quarantine and freed the board. The main run script always released clean. That script’s default path also never looked up the board’s address, so every run would have quarantined the board without reaching it, and the same clean release would have hidden that. Both were fixed.

Second, the design still quarantined on outcomes that were safe: a script on the board that refused and changed nothing, a command that ran to the end and failed, a capture that timed out, and a reserved lease the broker found after a restart. None meant the board was mid-change. The restart case quarantined because the broker couldn’t confirm the lease’s owner was alive, and a configuration switch can trigger it. I hit it myself, on a read-only check. I objected that if every bad run meant I had to clear a quarantine, the design put me back at the bench. I said the broker could clear one itself once a reboot worked, and asked for the plan to be reviewed for other steps that would call me back.

I told a session to build the revised plan, and the next commit to the broker narrowed the rule. A remote command that runs to completion is a known outcome, whatever it returns. Only a lost connection with no output, or a command that timed out, is uncertain. A refusal returns the broker to the state it was in before the operation. A capture timeout no longer counts as uncertain, because capturing doesn’t change the board. A restart reclaims a lease that changed nothing. Release is judged from the broker’s own ledger of operations, not from what the caller says. A caller can still force a quarantine when it releases, but a clean release can never clear one.

Third, a caller could still force a quarantine, and the first driver trial through the broker ended that way. Its one reboot came back. Then a check on the fresh boot refused, harmlessly, because the marker the loader leaves once the new driver loads was missing, most likely from a loader race that a later offline test reproduced. The script running the trial counted the unconfirmed trial as hardware doubt and forced a quarantine when it released. I authorized a clear, and after a loader fix and a capture fix, a manual run through the broker released the board clean. A caller that can force a quarantine carries its own copy of the rule, so narrowing the rule means narrowing it there too.

Send at most one guarded reboot per boot

Rebooting is the obvious way out of an unknown state, but a reboot can leave the board unable to answer at all. The reboot rules come from an outage before the broker existed. A CI job soft-rebooted the board while it ran an older NixOS image on the vendor’s kernel. That image had no reliable restart path. Only systemd’s runtime watchdog was armed, not its reboot watchdog, so a reboot that stalled had nothing to reset the SoC. The board never came back, and it sat unreachable for two days until I power-cycled it. CI kept running hardware jobs against the dead board, and main went red.

The first fix only contained the damage: a switch file committed to the board’s repository made those hardware steps skip with a warning, so a hung board no longer turns main red. The switch is on again because the hardware workflows target that older image, while the board now runs a mainline kernel. Pull-request CI checks software only. Hardware evidence comes from runs that hold a lease, and those leases have gone to workers and to me.

One tracker issue runs as a nine-step closed test loop.Racetrack loop, clockwise; the normal path needs no human step. 1 A worker is launched from a ready issue in the tracker. 2 The worker writes the patch on its issue branch. 3 The driver module pair is built elsewhere against the exact cached kernel. Steps 4 to 8 run inside the broker lease, which covers every step that touches the board, and produce the hardware evidence: 4 the broker stages files only over SSH; 5 guarded soft reboot, confirmed by a new boot id within the readiness window; 6 one traced driver start, expecting MAC_EN=0x4; 7 HDMI capture through the capture device, 2 s warm-up then a strict 15 s window; 8 diff of the 51 link-register words against the reference card, archived as evidence. 9 PR, CI and auto-merge: CI runs with the offline switch on, so it checks software only; the merge closes the issue and releases its dependents, and the loop returns to step 1. Invariant: one guarded reboot and one traced start per boot. Side exit: an uncertain broker outcome quarantines the board; the human_review recovery issue that would then ask the operator for a hand has not run live.broker lease ·hardware evidence1issue dispatchedfrom the tracker2patch writtenissue branch3module buildbuilt elsewhere4staged over SSHlease · files only5guarded rebootnew boot id6one traced startMAC_EN=0x47HDMI capture2 s warm-up · 15 s8diff + evidencevs reference · n/519PR, CI, auto-mergeoffline switch onCI: software checks onlyissue closes, deps releaseone reboot, one start per bootuncertain outcomequarantinedhuman_reviewrecovery issuenot run liveasks for a handoperator
One tracker issue as a loop: steps 4 to 8 run under the worker’s lease and yield the hardware evidence, and the guarded reboot in step 5 runs only after checks on the board pass. An unknown outcome quarantines the board, and the recovery path that would then file an issue asking me for a hand has never run live.

I still wanted a soft reboot wherever the workflow needed one, rather than unplugging cables for every restart. The broker allows them only under two rules. First, a table of the board’s installed images records each one’s SSH host key, which the broker checks to confirm which image it reached. The table also records whether an image’s reboots reset the SoC through the hardware watchdog, as the mainline image’s do without relying on systemd’s reboot watchdog. The broker refuses a soft reboot on any image the table doesn’t mark that way, or doesn’t list. Second, it sends at most one guarded reboot per boot.

A guard on the board checks the expected boot id, kernel, the exact NixOS system it booted, an idle display pipeline and the watchdog restart provider, then leaves a once-per-boot marker and schedules the reboot. The broker then waits up to ten minutes for the board to pass the same checks on a new boot id, which I’ll call a verified idle boot. If none appears, it never sends a second reboot: the outcome counts as uncertain, the broker records it and quarantines the board, and any physical recovery is mine. Two recorded boots took about four and a half minutes to reach SSH, and an earlier four-minute window once ran out while the board was still booting. A wait that ends early makes a slow boot look like a hung one.

Add workers once the broker can refuse a second holder

Because many tasks never touch the board, I asked for a second worker, but only once the broker could refuse a second holder on its own. Software-only work could then run beside a live measurement, and a second hardware task would wait or be refused.

No worker is allowed to clear a quarantine. That rule lives in the instructions each worker starts with, not in a check in the broker. So it is the kind of cooperative rule the broker replaced, and one more gap. A clear takes my say-so, which the broker records, or the broker’s own recovery, which may reboot the board once and clears the quarantine only on a verified idle boot. If it can’t get one, it files an issue in my tracker asking me for a hand. That path is built, held off by a flag, and has never run live. Clearing only on that boot means recovery accepts no state the normal loop wouldn’t.

The broker is the only service on the controller given the board’s login, and the gaps above are still open. When several kinds of caller share one scarce device, put its credential in one service’s configuration and nowhere else, and close every other way in. A lease, fencing tokens and more workers only mean something after that.

Share