nvidia-peermem sidecar crashloops after "userspace-only install", wedging RDMA nodes and the driver upgrade flow
Maintainers usually reply within 1 day
@kvalliyurnatt is already working on this.
Since Sep 28, 2026.
Assessment
This issue has not been assessed yet.
Description
Environment
- gpu-operator v26.3.2, driver 580.126.20 (open kernel modules), container-toolkit v1.19.1
- AKS, Kubernetes 1.35.5, Ubuntu 24.04, kernel 6.8.0-1059-azure, containerd 2.3.2
- Standard_ND96isr_H100_v5 nodes, driver.rdma.enabled=true (MOFED via network-operator DOCA driver container, useHostMofed=false)
Symptom
After a driver daemonset pod restarts on a node where the driver is already loaded, the pod sticks at 2/3 indefinitely:
nvidia-driver-ctr: ready=true (log: "performing userspace-only install")
nvidia-gdrcopy-ctr: ready=true
nvidia-peermem-ctr: CrashLoopBackOff
modprobe: FATAL: Module nvidia-peermem not found in directory /lib/modules/6.8.0-1059-azure
The daemonset never goes Ready, so the driver-upgrade state machine stalls (node stays cordoned in pod-restart-required / validation-required) and ClusterPolicy stays notReady.
Root cause
When the driver container starts and finds the desired driver version already loaded, it takes the reuse path: "The NVIDIA driver is already loaded with the desired configuration, performing userspace-only install" (nvidia-installer runs with --no-kernel-modules). That path never populates /lib/modules/<kernel> with the kmod tree in the new pod's mount namespace, and the container's startup cleanup removes any tree left by the previous pod. The nvidia-peermem sidecar then runs a disk-based modprobe, which fails — even though the module is loaded and fully functional in the kernel:
# lsmod (on the same node, at the same moment)
nvidia_peermem 16384 0
ib_uverbs 196608 67 nvidia_peermem,rdma_ucm,mlx5_ib
# find /lib/modules -name 'nvidia-peermem*' -> (nothing)
Restarting the pod can never recover: every restart re-detects the loaded driver and repeats the userspace-only install. The sidecar's check conflates "module file on disk" with "module available", and the reuse path breaks that assumption.
Two trigger paths observed (same day, same cluster)
- Manual driver pod restart (kubectl delete pod) with modules left loaded.
- The operator's own upgrade-controller roll (pod-restart-required): k8s-driver-manager's unload failed silently because a crashlooping dcgm-exporter pod still held /dev/nvidia-uvm handles; the flow proceeded anyway, landing the new pod on the reuse path. So this is reachable inside the fully managed path, not just via manual intervention.
Verification (retained evidence, post-recovery)
- Affected driver pod showed nvidia-driver-ctr restartCount=1 and nvidia-peermem-ctr restartCount=2; previous-container logs retain both the "userspace-only install" line (with the --no-kernel-modules warning) and the peermem load failure.
- Cluster events retain the peermem BackOff during the GPUDriverUpgrade state flow, plus dcgm-exporter restart/sandbox failures while the node cycled through pod-restart/validation.
- After recovery via full reinstall, the same nodes are healthy: driver pods 3/3, ClusterPolicy ready, nvidia_peermem loaded, and module files present under /lib/modules and /run/nvidia/driver/lib/modules.
Workaround (validated)
Per affected node, hold the daemonset off while unloading, then let it reinstall fresh:
kubectl label node <node> nvidia.com/gpu.deploy.driver=false nvidia.com/gpu.deploy.dcgm-exporter=false --overwrite
# wait for the driver + exporter pods to fully terminate (they hold /dev/nvidia*)
rmmod nvidia_peermem gdrdrv nvidia_uvm nvidia_modeset nvidia # on the node; retry if "in use"
kubectl label node <node> nvidia.com/gpu.deploy.driver=true nvidia.com/gpu.deploy.dcgm-exporter=true --overwrite
The label flip matters: with plain pod deletion, the recreated pod races the unload, re-detects the loaded driver, and re-acquires the devices. A node reboot is an equivalent, blunter alternative.
Suggested fixes
- The peermem sidecar should treat "module already loaded" as success (check /sys/module or lsmod before the disk modprobe), and/or
- The userspace-only install path should still lay down the kernel-module tree so disk-based modprobe stays valid, and
- k8s-driver-manager should fail hard (retry/backoff) when module unload fails, instead of proceeding into a state the sidecars can't handle.
- Dominant language
- Go
- Stars
- 2.9k
- Forks
- 569
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 76
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/gpu-operator
-
[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecarPossibly taken A pull request linked to this issue is open or already merged. Openbug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/gpu-operator#3026 ·
Maintainers usually reply within 1 day
-
[Bug]: GPUCluster common name label breaks DRA validator selectorPossibly taken @ajavanma claimed this 17 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2955 · 1 comment ·
Maintainers usually reply within 1 day
-
feature lifecycle/frozen needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 25/100
NVIDIA/gpu-operator#3035 · 1 reaction ·
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/gpu-operator#3027 ·
Maintainers usually reply within 1 day
-
Ensure automated backport commits have verified signaturesPossibly taken @asivanadi0 claimed this 8 days ago. Opengood-first-issue
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/gpu-operator#2997 · 1 comment ·
Maintainers usually reply within 1 day
All issues in NVIDIA/gpu-operator
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
open-telemetry/opentelemetry-go-compile-instrumentation#1467 ·
Maintainers usually reply within 3 days
-
Python 3.15 supportPossibly taken @amnesiaof claimed this today. OpenL: python L: python:uv
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
dependabot/dependabot-core#16524 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
duplication
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
openvibely/openvibely#1443 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 60/100
canonical/service-mesh#845 ·
Maintainers usually reply within 1 day