Cm5 macb network hang

From dtype.org

Silent network hang on Raspberry Pi CM5 / Pi 5 (macb / RP1 ethernet)

A technical writeup of a silent ethernet TX stall on the Raspberry Pi Compute Module 5, what it looks like, how to tell it apart from superficially similar bugs, and which mitigations did and did not work in our testing. Written for others hitting the same failure. Claims here are limited to what we directly observed or can cite; inference and open questions are marked as such.

Update (2026-10-04)

  • Three months on the patched driver, nothing has escalated. All five CM5s have run the backported driver (mitigation 4) without a reboot since early-to-mid July 2026: uptimes of 83–95 days, about 450 node-days combined. Three of the five also run the user-space reachability watchdog (mitigation 1), and it has not had to act once. Its last interface bounce was on 2026-06-29, before the patched driver was made persistent. The other two run patch 3 alone and have not needed a reboot either. Patch 3 logs only the first stall per boot, so we cannot count the stalls it caught. What we can say is that none of them needed anything beyond patch 3.
  • Ubuntu is about to ship mainline's timeout callback, and only that. linux-raspi 7.0.0-1021.21, in resolute-proposed since 2026-09-21, contains e438ec3e9e95 and net: macb: drop in-flight Tx SKBs on close. Both came in through Ubuntu's routine upstream-stable patchset (LP #2164666), not through bug #2133877, which is still "Confirmed". The build currently in -updates, 7.0.0-1020.20, has neither. Once 1021 or later is promoted, a stock Ubuntu Pi 5 / CM5 gets dev_watchdog-driven recovery and none of #7340. See State of the fix in Ubuntu.
  • That stock recovery has been seen failing to recover. In July 2026 the Raspberry Pi kernel (rpi-6.18.y) reverted #7340's patches 2 and 3 and took e438ec3e9e95 in their place. Patch 1 stays, now gated behind a capability flag. On that kernel, raspberrypi/linux #7661 (Pi 5, 6.18.50) shows the timeout firing and recovering every few hours for about a day. Then came a stall where it re-fired every 5 s for 3.4 hours without clearing. A candidate fix is in testing (PR #7667: reap lost TX completions on timeout, and fix BQL accounting in the TX error path). We have not run e438ec3e9e95 ourselves, so this is not a direct comparison with patch 3. It is still why we do not treat the stock kernel as a replacement for patch 3 plus a reachability watchdog. When 1021+ reaches -updates, we plan to rebuild the out-of-tree module against it rather than go stock.
  • Our TSO/GSO-off result does not cover "SG off". On #2133877 another reporter, also on Ubuntu linux-raspi 7.0 but stock (no #7340), found that ethtool -K eth0 tso off sg off stopped the stall: 0 stalls in about 256 node-hours, against about 18 expected from their baseline (comment #52; earlier comments #34 and #45 report the same). Our bundle (mitigation 3) turned off TSO and GSO but left scatter-gather on, and it hung. The two results do not conflict, and we have not tested SG off. The same comment reports the performance governor still stalling (2 stalls in 74 h), which agrees with our single trial (mitigation 2).

Update (2026-09-18)

This page was first written on 2026-06-29, while the fix below was still being tested. Since then:

  • The main fix patches did not prevent the stall. We ran all three #7340 patches on five CM5s. Patch 1 (PCIe posted-write flush) and patch 2 (ISR re-check) did not stop it: we saw at least two stalls in service, one about 8 hours after boot and one after 32 days (~10.9M TX descriptors). Patch 3, the in-driver watchdog, detected and recovered both, with about 1.7 s of lost egress in the case we timed. See mitigation 4.
  • Mainline took only a TX-timeout callback (e438ec3e9e95, "net: macb: add TX stall timeout callback to recover from lost TSTART write"). Patches 1 and 2 are not in mainline.
  • The stall is not specific to RP1 or PCIe. On netdev, Théo Lebrun (Bootlin) reproduced it on a Mobileye EyeQ5 Cadence GEM (rev 0x00070200) on current net/main, using iperf3 -c $IP -P10 -t3000. In his tests -P10 was required, even though one stream already reaches line rate. He suspects a race in macb_start_xmit; the root cause is still open.
  • Correction: an earlier version of this page said the kernel's dev_watchdog does not fire because trans_start keeps being updated. That was wrong. netdev_watchdog_up() only arms the watchdog when the driver implements .ndo_tx_timeout, and macb did not before e438ec3e9e95. So on older kernels (including Ubuntu's 7.0.0-1011) the watchdog never runs at all. With that commit it does fire on this stall in Bootlin's reproduction (NETDEV WATCHDOG: … transmit queue 0 timed out …).

Summary

On Raspberry Pi 5 / CM5 hardware (Cadence GEM MAC, macb driver, RP1 southbridge), the on-board gigabit ethernet can stop transmitting while the link layer continues to report the link as up. The receive path keeps working, the driver logs nothing, all ethtool error counters stay at zero, and on kernels without e438ec3e9e95 nothing in the kernel notices. The host remains fully alive locally; it is simply unreachable on the network until the NIC is reset (for us a link down/up always sufficed; at least one other reporter needed a driver unbind/bind).

We observed this on four of our five CM5s running Ubuntu 26.04 with the linux-raspi 7.0.0-1011 kernel. Based on a captured counter time-series (below), we attribute it to the silent TX-ring stall tracked in raspberrypi/linux #7339, with candidate fixes in PR #7340. None of

  1. 7340 is in Ubuntu's linux-raspi. Mainline's narrower timeout callback is

in -proposed as of 2026-10-04 (see Update), while Launchpad #2133877 is still "Confirmed". What we run now is the #7340 driver backported onto the Ubuntu kernel. It does not prevent the stall, but its in-driver watchdog recovers it in about a second or two.

Environment where we observed it

Item Value
Board Raspberry Pi Compute Module 5 (CM5), arm64
OS / kernel Ubuntu 26.04 LTS, linux-raspi 7.0.0-1011-raspi
NIC (from dmesg) macb 1f00100000.ethernet eth0: Cadence GEM rev 0x00070109, PHY Broadcom BCM54213PE
Topology On Pi 5 / CM5 the ethernet MAC is reached over PCIe via the RP1 southbridge (unlike Pi 4, where the MAC is on the SoC). We originally suspected this PCIe path. The stall has since been reproduced on a non-RP1 GEM (see Update), so it is not required.
Root filesystem NVMe (not SD). Relevant only because it makes the driver swap below low-risk.

We did not observe the failure on an x86-64 node with a different NIC in the same cluster, consistent with this being specific to the macb driver rather than anything in the surrounding software.

Symptom and signature

The defining property is that carrier never drops. A cable, switch, or PHY fault produces a Link is Down event; this failure does not.

Observed, every time:

  • ip link shows the interface LOWER_UP; the last carrier event in the log is the Link is Up - 1Gbps/Full from boot. No carrier loss is ever logged.
  • All off-host traffic stops simultaneously — gateway, peers, NFS, DNS. ARP for the gateway goes INCOMPLETE.
  • The macb driver emits nothing — no TX-timeout, no DMA/IRQ error. ethtool -S eth0 shows zero on every error and drop counter.
  • The kernel netdev TX watchdog (dev_watchdog) does not fire on kernels without e438ec3e9e95. That is because it is never armed (macb had no .ndo_tx_timeout), not because trans_start keeps ticking. On kernels that have the commit, it fires (see Update).
  • The host is otherwise healthy: the kernel is responsive on the local console, journald keeps writing, no panic, no OOM, no thermal event.
  • No self-recovery. The interface stays in this state indefinitely until reset.

If you see link-up + all-egress-dead + no carrier event + driver silent + zero error counters, you are very likely looking at this bug rather than a cable, PHY, congestion, or buffer-exhaustion problem.

Why it is easy to miss, and why we caught it quickly

This failure self-erases on reboot: the only evidence is in the previous boot's log, and that log shows nothing wrong with the link. On a desktop or a stateless worker it tends to read as a one-off "the machine dropped off the network," and a reboot makes it disappear. (Our first encounter, on an uninstrumented node, was exactly this — unreachable, no captured logs, reboot fixed it, cause unknown at the time.)

It became unmissable because the affected hosts were etcd voting members of a k3s control plane. When the NIC stalls on such a node, the local etcd is partitioned, the local API server can no longer serve quorum reads, the controller-manager loses its lease, and the k3s process exits and is restarted — a loud, logged, repeating failure rather than a silent drop. That is an artifact of our test setup, not of the bug, but it is a useful detail: if you run any quorum service (etcd, Consul, etc.) on CM5 hardware, this bug will surface as a control-plane meltdown, not as a quiet link blip. The underlying event is still just one stalled TX ring.

Diagnostic evidence we collected

Because none of the standard mechanisms detect this (carrier stays up, so systemd-networkd sees nothing; dev_watchdog is not armed for macb on our kernel; the hardware watchdog keeps being petted because the host is alive), we ran a reachability probe that, on N consecutive failures, captured a read-only forensic snapshot before resetting the interface. The snapshots are what let us classify the failure rather than guess at it.

What the snapshots showed, consistently:

  • Link up, 1000/Full, Link detected: yes at the moment of the hang.
  • Every error and drop counter at 0 (rx_resource_errors, rx_overruns, FCS errors, tx_carrier_sense_errors, q0_{tx,rx}_dropped). Rules out cable, congestion, buffer exhaustion, and PHY faults.
  • All four EEE/LPI counters ({rx,tx}_lpi_{transitions,time}) at 0 — the PHY had not entered Low Power Idle.
  • No RCU stalls in the kernel log preceding any hang.
  • A single TX queue (only q0_* counters exist) with the eth0 IRQ servicing entirely on CPU0.

The decisive capture was a 3-sample, ~1-second-apart time-series of the tx/rx frame counters taken at the moment of one hang:

Sample tx_frames rx_frames
t0 23553102 26633386
t1 23553102 26633400
t2 23553102 26633423

tx_frames is frozen across all three samples while rx_frames continues to climb, and the eth0 IRQ count keeps advancing. This is a direct observation of the TX path being stalled while RX and interrupts are still live — not an inference from symptoms. It is this datapoint, more than the symptom description, that drives the root-cause attribution below.

Root-cause analysis: which bug is this?

There are several public reports that look alike at the symptom level. They are not all the same bug, and getting the mechanism right determines which mitigation is worth applying. This section lays out the candidates and the evidence for our attribution.

The candidate reports

Report Proposed mechanism Same silicon?
Ubuntu Launchpad #2133877 — "Complete network hang on Raspberry Pi 5 … possibly related to CPU frequency scaling" CPU frequency-scaling transitions corrupting the RP1/macb DMA path; the report notes RCU stalls as a precursor and that the performance governor stopped those RCU stalls. Yes (Pi 5, linux-raspi)
raspberrypi/linux #7339 / PR #7340 — "candidate fixes for silent TX stall on BCM2712/RP1" A silent TX stall in the macb driver: tx stops advancing, RX keeps working, link stays up. Three driver patches. Yes (BCM2712/RP1, macb)
siderolabs/sbc-raspberrypi #91 — "silent network death on Talos" Decomposes the failure into an EEE LPI-wake race, a macb TSO/GSO TX-ring hang (single TX queue, softirqs on CPU0, small rings), and a Talos-specific socket issue. Prevention bundle: EEE off + TSO/GSO off + larger rings. Yes (Pi 5, BCM2712 + Cadence GEM + BCM54213PE)
Various "ethtool -K eth0 tso off gso off fixes it" posts Trace originally to Intel e1000e "Detected Hardware Unit Hang" — a different NIC and driver. No

What our evidence supports, and what it argues against

The symptom matches #2133877 essentially line-for-line — same NIC, complete network hang, host alive locally, clean ethtool stats, no macb log lines, recovers only on reset. So at the symptom level we are clearly in the

  1. 2133877 family.

But #2133877's proposed mechanism (frequency scaling) does not fit our observations, on three independent points:

  1. No RCU stalls. #2133877's headline precursor is RCU stalls (the reporter saw them at a steady rate, and pinning the governor stopped them). We logged zero RCU stalls before any hang, across every event. The precursor that report hangs its causal story on is absent on our hardware.
  2. The hang occurred at maximum, pinned frequency. We trialed the performance governor (which holds the cores at a fixed maximum frequency and therefore eliminates frequency transitions) on the worst-affected node. It hung again ~30 minutes later, and our counter capture from a separate hang shows the cores already at the maximum 2400 MHz at the time of the stall. If frequency transitions were necessary to trigger the hang, pinning the frequency should have prevented it. It did not. (This is a single negative trial, n=1, but it is a direct one.)
  3. The captured time-series matches #7340's description exactly. #7340 describes the failure as "tx_packets stops incrementing, RX still works, link stays up." Our capture shows precisely that — frozen tx_frames, climbing rx_frames, live IRQ. #2133877 does not characterize the failure at the ring-counter level; #7340 does, and our data fits it.

Our interpretation: #2133877 and #7339 are most likely the same underlying defect, and the "frequency scaling" wording in #2133877's title is a hypothesis from before the TX-ring stall was isolated, not an established mechanism. The work that isolated the TX-ring stall and produced candidate patches is #7339/#7340, whose described signature is the one we captured. We treat the frequency-scaling angle as not supported on our hardware: we cannot rule out a frequency transition acting as one possible trigger, but our evidence shows transitions are not necessary (the hang happens at pinned max frequency) and that the RCU-stall precursor central to that report is absent for us.

On the EEE and TSO/GSO theories (siderolabs #91)

siderolabs #91 is the closest same-silicon analysis and is the source of the popular prevention bundle (EEE off, TSO/GSO off, larger rings). Two of its three triggers are worth separating against our data:

  • EEE / LPI-wake race — not our mechanism. All four LPI counters read zero at every hang; the PHY never entered Low Power Idle. Disabling EEE addresses a state our hardware was never in. (On our setup EEE was advertised but inactive because the switch did not negotiate it, so disabling it also had no measurable power cost — but that is incidental.)
  • macb TSO/GSO TX-ring hang — this is the trigger that does line up structurally with what we see: a single TX queue and all NET softirqs on CPU0. (Its explanation of the silence, that trans_start keeps ticking, is wrong for the reason given in the Update.) This is the same family as #7340. However — see the mitigations section — disabling TSO/GSO and enlarging the rings did not prevent the hang on our kernel, so while the structural description fits, the offload-disable remedy did not hold for us.

The e1000e-derived "just turn off TSO" advice does not transfer — it originates from a different NIC and driver, and a lone TSO-off was tried and walked back in the

  1. 7340 discussion. We mention it only because it is widely repeated.

What we are NOT claiming

  • We have not bisected the kernel or independently tested the PCIe posted-write mechanism behind patch 1. Two things now argue against it: patch 1 did not prevent the stall for us, and the stall reproduces on non-PCIe GEM hardware.
  • The backported fix does not eliminate the stall (see mitigation 4). We cannot measure how often it still stalls, because patch 3 logs only the first stall per boot.
  • We cannot speak to whether other Pi 5 / CM5 configurations (different PHY link speeds, RPiOS vs Ubuntu, SD vs NVMe) behave identically.

The upstream fix

raspberrypi/linux PR #7340 was merged into the Raspberry Pi Foundation kernel branch rpi-6.18.y on 2026-05-08, and the series was then posted to mainline netdev (v2, 2026-05-14). It is three patches to the Cadence/macb driver. Mainline took none of them as-is. It merged only a TX-timeout callback, e438ec3e9e95 (2026-06), which calls the existing macb_tx_restart() from .ndo_tx_timeout.

# upstream commit Change (per the commit) Role
1 63d230184da3 Flush the PCIe posted write after the TSTART doorbell, so the doorbell that tells the MAC to begin transmission is not left sitting in the PCIe fabric. Proposed as the core fix. Not merged in mainline (doubted in review); did not prevent the stall for us
2 3ccf780ce058 Re-check the interrupt status register after re-enabling interrupts in macb_tx_poll. Closes a lost-interrupt race. Not in mainline; did not prevent the stall for us
3 8ea87c96f1a6 Add a TX-stall watchdog inside the driver: a per-queue 1 s delayed_work that re-kicks TSTART when tx_tail has not moved on a non-empty ring. Defence-in-depth. This is what recovers the stall for us. Mainline's e438ec3e9e95 does the same job via dev_watchdog, which only trips once the TX queue has stopped (ring full or BQL limit) and trans_start is ≥5 s old. The RPi kernel replaced this patch with e438ec3e9e95 in July 2026 (see Update, 2026-10-04)

Patch 1 implies a mechanism: a posted MMIO write delayed in the PCIe fabric, so the MAC never sees the "start TX" doorbell. That fits the symptoms, but we no longer think it is the cause. Patch 1 did not prevent the stall in our deployment, and the stall reproduces on a GEM that has no PCIe hop (see Update).

Caveats from the upstream discussion (worth knowing if you deploy the fix):

  • Upstream labels them "candidate fixes." Contributors reported that a software watchdog could still occasionally trigger even with the patches applied, and that "both patches + watchdog are necessary to keep things stable."
  • Recovery experience varied: a ip link down/up recovered the interface for some reporters but not all — at least one needed a driver bind/unbind. On our nodes a bounce cleared every one of the 9 instrumented hangs; none needed the reboot fallback.

State of the fix in Ubuntu

As of 2026-10-04:

Build Pocket macb changes since our 7.0.0-1011
7.0.0-1020.20 resolute-updates / -security none
7.0.0-1021.21 resolute-proposed (since 2026-09-21) e438ec3e9e95 (TX stall timeout callback) and drop in-flight Tx SKBs on close, both through upstream stable patchset 2026-08-20 (LP #2164666)

Because mainline took only the timeout callback, an Ubuntu update can only bring e438ec3e9e95, not the #7340 patches. The fix is not arriving through Launchpad bug #2133877. That bug is still "Confirmed", and a reporter there asked on 2026-09-07 for e438ec3e9e95 to be cherry-picked. So the bug's state is not a reliable signal; watch the package instead. To see which pocket a build is in: https://api.launchpad.net/1.0/ubuntu/+archive/primary?ws.op=getPublishedSources&source_name=linux-raspi&exact_match=true&status=Published.

Checking for it: an earlier version of this page suggested apt-get changelog linux-image-raspi | grep -iE 'TX stall|2133877|TSTART'. That check is unreliable. The meta package's changelog is the wrong one, Ubuntu's changelog lines are upstream commit subjects that never carry the bug number, and "TX stall" also matches unrelated drivers. Instead, fetch the versioned source changelog (https://changelogs.ubuntu.com/changelogs/pool/main/l/linux-raspi/), keep only the entries newer than your installed version, and grep that part for macb and for "TX stall timeout callback".

Mitigations and their measured results

We tried four things; two are worth running. The backported driver's in-driver watchdog (patch 3) is what now recovers the stall. A user-space reachability watchdog sits behind it as a last resort. Nothing we tried prevents the stall. We report what each one actually did, including the ones that failed.

1. Reachability watchdog (recovery) — works; now the last-resort layer

A small periodic probe that pings two independent off-host targets (we use the gateway and a NAS). Requiring both to fail in a cycle, and requiring N consecutive failed cycles, avoids false positives from ordinary blips. On trip it:

  1. Captures a read-only forensic snapshot (the data that made this writeup possible).
  2. Resets the interface: ip link set eth0 down && ip link set eth0 up. This re-initialises the MAC; in our 9 instrumented hangs it cleared the stall every time, without a reboot.
  3. If still dead after the bounce, reboots as a last resort.

Notes that matter:

  • It must test reachability, not link/carrier state — carrier stays up, so a link-state monitor will never fire.
  • The recovery script and its logs must live on local disk, not on any network filesystem — the network is exactly what is gone when it runs.
  • Timing has to beat whatever your stack's failure threshold is. In our case a slow bounce let the node stay partitioned long enough for the k3s control plane to give up on etcd and restart. We tightened detection (5 s probe cadence, bounce after 2 failed cycles ≈ ~10–15 s to first reset) while keeping the reboot a patient last resort (~75 s). If you run quorum services, tune the bounce to land inside their lease/election timeout.
  • A reboot is an acceptable recovery for a node that is one member of a fault-tolerant quorum; it is more disruptive for a singleton. The bounce-first ordering exists to avoid rebooting when a link reset would do.

The on-board hardware watchdog cannot substitute for this, because the host itself is not hung. On kernels without e438ec3e9e95, dev_watchdog is not armed for macb either.

Before the patched driver, this watchdog was the only recovery. With the patched driver in place it has not had to act once (through 2026-10-04, about 280 node-days on the three nodes that run it). We keep it running in case patch 3 ever fails.

2. performance CPU governor — did NOT prevent it

Tested specifically to evaluate the frequency-scaling hypothesis. Pinning the cores to maximum frequency eliminates frequency transitions. The node hung again ~30 minutes after pinning. We reverted it. Besides being ineffective here, it carries a small but continuous idle power cost (we measured roughly +0.5–0.75 W on a fanless ~3 W-class node). Recommendation: do not rely on this, and treat it as evidence against the frequency-transition mechanism rather than a mitigation.

3. EEE / TSO / GSO off + larger rings — did NOT prevent it on our kernel

This is the siderolabs #91 prevention bundle:

<syntaxhighlight lang="bash"> ethtool --set-eee eth0 eee off ; ethtool --set-eee eth0 advertise 0x0 ethtool -K eth0 tso off gso off ethtool -G eth0 rx 4096 tx 2048 </syntaxhighlight>

We applied and verified all of it (confirmed active at the time of a subsequent hang: EEE disabled, both offloads off, rings 4096/2048). The node hung anyway. On our kernel this bundle did not prevent the stall, so we cannot report it as effective prevention. A fleet reporter on other kernels (notably pre-#7340 builds) found it stable, so do not assume their result carries over to linux-raspi 7.0. We removed it on 2026-08-02: TSO/GSO off moves segmentation onto CPU0 and bought us nothing. Since then all nodes run stock NIC settings with the patched driver, and no stall has needed more than patch 3 to recover.

Note that we left scatter-gather on. A reporter on #2133877 found ethtool -K eth0 tso off sg off stopped the stall on a stock 7.0 kernel (see Update, 2026-10-04). We have not tested that variant, so our negative result does not apply to it.

If you do apply it, note that a link down/up resets these settings, so any watchdog that bounces the interface must re-assert them afterward.

4. Backported macb.ko — the running state; patch 3 is what recovers it

Because the #7340 patches are driver changes and macb is a loadable module (CONFIG_MACB=m) on the Ubuntu kernel, the three #7340 patches can be backported and the driver rebuilt without replacing the kernel. The patches applied cleanly to the linux-raspi 7.0 source. Patches 1–2 were aimed at the root cause; patch 3 is recovery.

We first loaded it live and non-persistently, inserted at runtime from outside /lib/modules, so any reboot returned to the stock driver. With root on NVMe (so the NIC driver is not boot-critical) that is low-risk. We made it persistent on 2026-07-01, and the two-week soak finished with no hangs. The module goes in /lib/modules/$(uname -r)/updates/, which overrides the in-tree driver, the initramfs is regenerated, and the kernel package is held (apt-mark hold linux-image-raspi). A loader script did the live swap (verify the module's vermagic matches the running kernel, rmmod stock → insmod patched → bring the link up → wait for a real off-host ping → roll back to stock on any failure), with the watchdog as the final backstop.

Two operational warnings if you try this:

  • A live macb reload destroys and recreates eth0. If you run an overlay network (we run flannel/VXLAN), the overlay device is bound to the old interface and may not rebuild itself — leaving the host with full LAN connectivity but no overlay (cross-node traffic silently fails). The fix is to restart the networking layer that owns the overlay after the swap. This bites both the swap and a manual rmmod; modprobe revert.
  • The module is bound to one exact kernel version (vermagic). Any kernel upgrade obsoletes it, and it must be rebuilt. A reboot after an upgrade but before a rebuild simply runs the stock driver: safe, just unprotected. We hold the kernel and persist the exact module we tested. If you don't hold the kernel, use DKMS, which rebuilds the module on each kernel bump.

The module loads tainted (out-of-tree + unsigned, taint mask 12288), as expected.

Result

Before the patch, the three instrumented nodes hung 9 times in about 53 node-hours (1 per ~6 node-hours; the worst node, 1 per ~2.6 h). Each hang meant an unreachable host until reset. On the patched driver, the two-week soak finished with no hangs, and it now runs persistently on five CM5s. At least two in-service stalls have occurred since (one ~8 h after boot, one ~32 days / ~10.9M TX descriptors in). Patch 3 caught both and logged:

<syntaxhighlight lang="text"> macb 1f00100000.ethernet eth0: TX stall detected on queue 0 (tail=10917116 head=10917119); re-kicking TSTART </syntaxhighlight>

In the one we timed, egress was dead for about 1.7 s. There was no reboot and no user-space intervention. So the patches turn an unrecoverable hang into a one-to-two-second blip, and patch 3 does that, not patches 1–2. Two caveats:

  • Patch 3 uses netdev_warn_once, so it logs at most one stall per boot. The true stall rate is unknown, and we have not measured whether patches 1–2 lower it. Bootlin's reproducer (see Update) now makes that testable.
  • Some nodes log a tail=0 head=2 stall ~9 s after boot (two of our five). This is a start-up artefact, not a real stall, but because of netdev_warn_once it uses up that boot's only log line, so later real stalls on that boot go unlogged. Classify by time since boot: seconds = artefact, hours to days = real.

How to confirm you have this specific bug

After a recovery, examine the boot that failed (e.g. journalctl -b -1):

<syntaxhighlight lang="bash"> ethtool -i eth0 # driver should be 'macb' journalctl -b -1 | grep -iE 'Link is Down|carrier' # expect NOTHING for eth0 (link never dropped) journalctl -b -1 | grep -iE 'i/o timeout|not responding, timed out' # all egress dead at once journalctl -b -1 | grep -icE 'oom-kill|killed process' # expect 0 (rules out memory) journalctl -b -1 | grep -icE 'rcu.*stall' # we saw 0 (the #2133877 precursor) ethtool -S eth0 | grep -vE ': 0$' # at hang time, error/drop counters are 0 </syntaxhighlight>

If you can catch it live (before resetting the interface), the conclusive check is to sample the frame counters a second apart and confirm TX is frozen while RX moves:

<syntaxhighlight lang="bash"> for i in 1 2 3; do

 ethtool -S eth0 | grep -E 'tx_frames|rx_frames'; echo ---; sleep 1

done

  1. tx_frames identical across samples + rx_frames increasing = the silent TX stall

</syntaxhighlight>

On a patched driver the stall usually recovers before you can sample it. Check the kernel log instead: journalctl -kb | grep -iE 'TX stall|NETDEV WATCHDOG' (patch 3 and e438ec3e9e95 respectively).

Status and open questions (2026-10-04)

  • Confirmed: the failure is a silent TX stall (captured: TX frozen, RX live). It is not memory, thermal, PHY or cable, and it recovers on interface reset.
  • Confirmed not the cause for us: EEE/LPI (counters at zero); frequency transitions as a necessary trigger (it hangs at pinned max frequency); the RCU-stall precursor from #2133877 (never observed).
  • Did not prevent it: the performance governor; EEE/TSO/GSO off + larger rings (SG left on); #7340 patches 1–2.
  • Recovers it: #7340 patch 3 (in-driver, ~1–2 s); a reachability watchdog as the last resort (~10–15 s). About 450 node-days on patch 3 with no reboot and no action from the reachability watchdog, which runs on three of the five.
  • Reported elsewhere, untested by us: tso off sg off preventing it on stock Ubuntu 7.0 (#2133877); mainline's e438ec3e9e95 failing to clear a stall for 3.4 h on the RPi kernel (#7661).
  • Open: the actual root cause. It is reproducible off-RP1, and a driver race in macb_start_xmit is suspected (Bootlin, netdev, 2026-09). Also open: whether patches 1–2 lower the stall rate at all.
  • Upstream: mainline has only the e438ec3e9e95 timeout callback, and Ubuntu has it in -proposed (7.0.0-1021.21). As of 2026-10-04 no fix for the underlying race has been posted to netdev.

References