· · 12 min read

Finding a DSI driver bug by diffing registers against the vendor image

A register diff is only as good as the blocks its baseline covers. A review of a DSI start failure ruled out every candidate register fix, yet the fault sat in a block its baseline skipped, where a masked write had left an undocumented bit set.

Diffing registers means running a vendor’s software and your own on the same hardware and comparing the registers each leaves, word for word; every difference is a candidate cause. My board uses the SG2000 SoC, and a bridge chip turns the SoC’s MIPI DSI output, a serial display link, into HDMI. On a mainline kernel, a new driver for the SoC’s DSI host, the SoC-side hardware that drives the link, kept timing out when starting the link. A community Debian image on the vendor’s kernel and display drivers, which I’ll call the vendor image, gets that link into high-speed mode, which carries video. The new driver’s image and the vendor image run from separate SD cards in the same board, so a difference points at software, not silicon. The vendor image’s own output is a fine vertical pinstripe, not a picture, so matching it means matching link timing, not pixels. Diffing against it found differences that fixed nothing, and one comparison declared the registers matched while the link still timed out.

Use a reference that reproduces, and count boots

I made the vendor image’s SD card, the vendor card from here on, the reference, because it was the one setup whose output reproduced under measurement. I’ll call its pinstripe the stripe. A reference is measured, never developed on: once you change it, a difference can come from your change instead of from the new driver. Results count in boots, not trials: a trial is another start of the display on a running board, and on this board a second start on the same boot can behave differently from the first.

I had the card’s recipe, binaries and hashes saved rather than a whole-card image, which would have needed a second computer with room for it. Those files have never been tested as a reinstall, so one rewrite of that card could still lose the reference.

The alternative was the image already under test, a NixOS build on the vendor’s kernel for the ARM core (the SG2000 runs Linux on an ARM or a RISC-V core). Its output had gone flat, meaning the uniform dark frame the USB capture device on the HDMI output records when it gets no usable signal. Reflashing an older image wouldn’t have separated board from software, because the older image had also gone flat after restarts. An AI coding session, a session from here on, planned a short physical bisect, and I did the steps at the bench. After I cut the power, that image stayed flat. With the vendor card in and the core switch on RISC-V, the same board, cable and capture device showed the stripe. The bisect ruled out a damaged board. Leftover power state stayed possible.

The ARM image did draw the stripe in one frame after a core-switch round trip, but it didn’t hold up. Two first starts with the same recipe, each from a display block whose enables read 0 after a full power drain, gave one stripe and one flat frame. A session’s audit found that an earlier count of five successes was five trials on one boot. So I stopped chasing why it only sometimes drew the stripe. The new driver runs on the RISC-V core, like the vendor card. Drop a setup whose best result doesn’t repeat from a clean start, and report results by boot, so a streak on one boot never reads as a rate.

Know what the reference can prove

On one boot, three valid recordings from the vendor card gave 39 of 39 settled one-second samples that the capture tool classified as the stripe. Two recordings followed a restart of the display with the same recipe, and one only reopened the capture. That shows the stripe repeats on that boot, restarts included; counted by boots, it is one result. That boot’s first recording, which failed on a corrupt capture buffer, is kept as a failure. Recordings on three other boots were also classified as the stripe. The reference’s write-up lists what the recordings don’t establish: clean cold boots, a perfect first start, or arbitrary pixels on screen.

Parity with the vendor card means the DSI link enters high-speed mode and the bridge locks onto its timing. I haven’t found out why the vendor image draws a stripe, even when fed real images. Timing on the same silicon is still the right target, because timing is what the new driver lacked. Write down what a reference’s own output shows before you diff against it, because parity proves only that.

Three enlarged crops side by side: fine green vertical stripes, a flat near-black frame, and fine vertical stripes in olive and violet grays.
The same 64 by 36 pixel window of three captures, enlarged 8 times: the vendor card's own output, a flat frame from R14 (an earlier revision of the new driver), and R19 (the revision that cleared the timeout) on a warm boot. Colors differ, so only structure is compared; a match shows link timing, not a requested image.

Test each difference as a cause

From revision R8, the new driver’s start request to the DSI MAC, the host’s controller block, timed out every time it ran; earlier booted revisions had hung or failed in other ways. The request writes 0x4, the high-speed video bit, to the MAC’s enable register, which reads back 0x4 on the vendor card while its link streams. The driver polls that register for up to 20 ms, and it stayed 0. The timeout was a symptom: R14, which made it non-fatal, let the rest of the start run, but the enable register still read 0 and the capture stayed flat. A burst of 256 back-to-back reads of the enable register right after the request never saw 0x4 either, an observation at 256 instants rather than proof that high-speed mode never began. Each booted revision after R8 ruled out one explanation, from timing-generator order to settle time, without fixing the start.

I asked for a register comparison against the vendor card, the comparison from here on. Its first diff found that the new driver had turned off end-of-transmission (EoT) packets, which close each high-speed transmission. The mainline bridge driver asks for none, while the vendor image leaves the EoT bit set. I authorized one scoped write on the running reference: clear the EoT bit, capture, set it back. The stripe held both ways, so a running link doesn’t need EoT, and the new driver still timed out with EoT on. The difference was real, but it didn’t explain the timeout. A difference is a lead until one scoped test shows it’s a cause.

The D-PHY, the host’s physical layer, drives the lanes, the link’s clock and data pairs. From R15 the driver’s register snapshots, copies of a fixed range of each block’s registers taken at three points in a start, included the D-PHY. The snapshots showed the D-PHY word at offset 0x44 drop to 0 across the start, and a session’s report read that as the lanes leaving their low-power stop state. The driver had written both values itself. With both writes removed, the word stayed 0 and no D-PHY word in the snapshots moved. Before you read a register as a measurement, check that your own driver didn’t write it.

List the blocks your baseline covers

The header of R18, the last revision before the fix, summed up the comparison as one mismatch, the MAC’s enable register, which was the symptom itself, and declared register parity complete. An AI review of the vendor source, which worked from the comparison’s baseline and checked each candidate adversarially, refuted every candidate register fix, and its report blamed the analog hardware.

The header’s parity covered only the registers compared, on the pages read. The vendor card’s dump behind it, the comparison’s baseline, read the five pages on an existing dump script’s list: the DSI MAC plus the scaler, display timing, video subsystem control and clock controller. A full August dump of the vendor card’s MAC and D-PHY pages was already on disk:

BlockBaselineSnapshotsAugust dump
DSI MACyesyesyes
The other fouryesyesno
D-PHYnofrom R15yes

The comparison had no vendor page to hold the D-PHY snapshots against, and said so. Yet the review’s report put the fault in the D-PHY and called it invisible to registers. It was a register, in the one block the baseline didn’t cover. Comparisons on the ARM image had used the August dump, but the new driver’s baseline never picked it up.

The same gap had hidden a difference before. On the ARM image, a comparison called both stacks identical because every configuration register it read matched; it hadn’t read the data lanes’ state words or four bridge status bytes, which differed.

Before you trust “no differences,” write down which blocks, and which registers in them, each side of the diff read. A block missing on either side is unknown, not equal.

Port whole-register stores as whole-register writes

The review’s next step needed a view of the D-PHY: a scope on the DSI lanes or a supervised dump of the vendor card’s D-PHY page. I only had a multimeter. So I switched the board to the vendor card before I left and asked for a prompt so a session could prepare that dump.

An earlier userspace read of the D-PHY page had hung the vendor card, and this time nobody could get to the bench to recover it, so the automatic reviewer that screens a session’s commands before they run refused twice. The August dump had resurfaced while the live read was being vetted, which made that read a check on it. Rather than wait until I was back, I approved exactly one prepared read-only dump of the MAC and D-PHY pages, 128 words, 32-bit reads only, with no writes, no reboot and no other register windows. It came back with the card responsive and agreed with the August dump on all 51 D-PHY words the driver’s snapshots read; one word past that range differed between the two dumps.

Of those 51, 6 differed between the vendor card and every R18 snapshot; the August dump already showed all six, and the live read confirmed them. Four were lane-state readings, one an unnamed word, and one the power-down word at offset 0x64: 0 on the vendor card, 0x2000 on R18. Sort a diff’s words by who sets them: the lane states are status the hardware reports, and the power-down word was the only one of the six the driver wrote.

The vendor’s init does a whole-register store there, zero in every bit for this board’s lane setup. The port had turned it into a masked update, which reads the register, replaces only the bits in a mask and writes the word back with every other bit as it found it. Its mask, 0x1f1f, covers bits 0 to 4 and 8 to 12, so any other bit already set stayed set. The leftover 0x2000 is bit 13, one place above the mask. An earlier source comparison had called the port a register-for-register match for the vendor’s init, reading that masked clear as a store of zero. None of the sources the sessions inspected documents bit 13, and I don’t know what set it.

R19, which an AI session wrote, makes one change to what the driver does: it replaces the masked update (offset, mask, then the value for the masked bits) with a whole-register write.

-	sophgo_dphy_update(phy, 0x64, 0x1f1f, 0);
+	writel(0, phy + 0x64);

From the R19 patch, unedited.

R19 is a parity change rather than an understood mechanism, as the patch itself says. The alternative, explaining bit 13 first, had no documentation to start from. I asked for its trial on the board, and two checks came first. A session’s disassembly of the vendor module actually installed on the vendor card confirmed the whole-register store. That binary also differed from the vendor source in three other places, so the source alone couldn’t show what the card ran. And a host-side regression test, which runs the driver’s preparation code offline on simulated registers preloaded with 0x2000, failed on R18’s code and passed on R19’s. From a zeroed register, both would pass. It shows the new code clears the bit, not that the bit was the fault; only the board could show that.

When you port vendor init, reproduce a whole-register store as a whole-register write unless you know every bit you’re preserving.

Scope the result to what you measured

On one warm boot (a restart without cutting power) and one start of the display, R19 read the enable register as 0x4, the high-speed value the 256 reads never saw. In the snapshot taken once the link started, all 51 D-PHY words matched the vendor card, including the five differing words the driver never writes, and the timeout was gone. The capture recorded 882 frames at 60 frames per second in the reference’s stripe class (the figure’s right panel). Its first frame arrived later than the capture tool’s strict timing check allows, for a reason still unknown. Some of the vendor card’s own recordings did the same.

That’s link timing parity on one warm boot, and nothing more. It doesn’t show a clean cold start. On this board that needs HDMI unplugged and about 12 seconds of power drain, because cycling only USB power with HDMI connected kept the display block’s state. This is the leftover power state the bisect couldn’t rule out. It doesn’t show requested pixels reaching HDMI: a later run of the new driver that asked for solid red captured the same stripe. And I never put a scope on the lanes.

AI sessions wrote the patches, ran the source review and found the six differing words; I chose the reference, asked for the comparison, approved the one dump and committed R19 once it ran. A “no differences” verdict holds only for what both sides read, so the list of what they read belongs in the verdict.

Share