[Bug]: Crashlooping nvidia-peermem-ctr: could not insert 'nvidia_peermem': Invalid argument
Maintainers usually reply within 1 day
@kvalliyurnatt is already working on this.
Since Sep 9, 2026.
Assessment
This issue has not been assessed yet.
Description
Describe the bug
Filing a bug after the conversation at https://github.com/NVIDIA/gpu-operator/issues/2844 thanks
On a GB300 on-prem cluster we're seeing constant crashloops of nvidia-peermem-ctr:
DRIVER_ARCH is aarch64
modprobe: ERROR: could not insert 'nvidia_peermem': Invalid argument
failed to load nvidia-peermem module
Recent dmesg -T | grep -i -E 'peermem|ib_core|peer_mem':
[Tue Sep 8 17:16:02 2026] nvidia_peermem: disagrees about version of symbol ib_register_peer_memory_client
[Tue Sep 8 17:16:02 2026] nvidia_peermem: Unknown symbol ib_register_peer_memory_client (err -22)
[Tue Sep 8 17:16:02 2026] nvidia_peermem: disagrees about version of symbol ib_unregister_peer_memory_client
[Tue Sep 8 17:16:02 2026] nvidia_peermem: Unknown symbol ib_unregister_peer_memory_client (err -22)
[Tue Sep 8 17:21:18 2026] nvidia_peermem: disagrees about version of symbol ib_register_peer_memory_client
[Tue Sep 8 17:21:18 2026] nvidia_peermem: Unknown symbol ib_register_peer_memory_client (err -22)
[Tue Sep 8 17:21:18 2026] nvidia_peermem: disagrees about version of symbol ib_unregister_peer_memory_client
[Tue Sep 8 17:21:18 2026] nvidia_peermem: Unknown symbol ib_unregister_peer_memory_client (err -22)
lsmod | grep -E 'nvidia|mlx5|ib_':
mlx5_dpll 196608 0
mlx5_vdpa 196608 0
ib_ipoib 262144 0
ib_cm 327680 2 rdma_cm,ib_ipoib
ib_umad 262144 0
mlx5_fwctl 262144 0
fwctl 262144 1 mlx5_fwctl
mlx5_ib 655360 0
ib_uverbs 327680 2 rdma_ucm,mlx5_ib
mlx5_core 3145728 3 mlx5_dpll,mlx5_fwctl,mlx5_ib
mlxfw 262144 1 mlx5_core
mlxdevm 655360 1 mlx5_core
ib_core 720896 8 rdma_cm,ib_ipoib,iw_cm,ib_umad,rdma_ucm,ib_uverbs,mlx5_ib,ib_cm
mlx_compat 196608 15 mlx5_dpll,rdma_cm,ib_ipoib,mlxdevm,mlx5_fwctl,iw_cm,ib_umad,mlx5_vdpa,fwctl,ib_core,rdma_ucm,ib_uverbs,mlx5_ib,ib_cm,mlx5_core
nvidia_modeset 2097152 0
nvidia_uvm 1966080 8
nvidia 15138816 55 nvidia_uvm,gdrdrv,nvidia_modeset
video 262144 1 nvidia_modeset
ecc 196608 1 nvidia
nvidia_cspmu 196608 0
arm_cspmu_module 262144 1 nvidia_cspmu
macsec 262144 1 mlx5_ib
psample 262144 1 mlx5_core
tls 327680 1 mlx5_core
pci_hyperv_intf 196608 1 mlx5_core
To Reproduce
No clear way to get to those errors. Half of the nodes seem to load the module, half do not. No reasonable split of configs / kernels / ... so far explained this split.
Expected behavior
Not crash :-)
Environment (please provide the following information):
- GPU Operator Version: v26.7.0
- OS: Ubuntu24.04
- Kernel Version: 6.8.0-124-generic-64k or 6.8.0-136-generic-64k
- Container Runtime Version: containerd 2.2.3
- Kubernetes Distro and Version: k8s v1.35.3
DOCA: doca3.5.0-26.07-0.7.7.0-0-ubuntu24.04-arm64
@rajathagasthya LMK if a full debug bundle would be helpful
@lalitadithya FWIW: in this case nvsentinel is not actually picking up any errors, interestingly. The only "signal" from an administrator perspective is that pods are crashlooping continuously (which perhaps should be enough, but ... do consider if detection here would make sense thanks).
- Dominant language
- Go
- Stars
- 2.9k
- Forks
- 569
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 76
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/gpu-operator
-
[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecarPossibly taken A pull request linked to this issue is open or already merged. Openbug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/gpu-operator#3026 ·
Maintainers usually reply within 1 day
-
[Bug]: GPUCluster common name label breaks DRA validator selectorPossibly taken @ajavanma claimed this 17 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2955 · 1 comment ·
Maintainers usually reply within 1 day
-
feature lifecycle/frozen needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 25/100
NVIDIA/gpu-operator#3035 · 1 reaction ·
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/gpu-operator#3027 ·
Maintainers usually reply within 1 day
-
Ensure automated backport commits have verified signaturesPossibly taken @asivanadi0 claimed this 8 days ago. Opengood-first-issue
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/gpu-operator#2997 · 1 comment ·
Maintainers usually reply within 1 day
All issues in NVIDIA/gpu-operator
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
prime-radiant-inc/evener#4223 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
open-telemetry/opentelemetry-go-compile-instrumentation#1467 ·
Maintainers usually reply within 3 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
yetone/magpie#1490 · 1 comment ·
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
a:bug
Difficulty 2/5 1-3 hours Newbie friendliness 80/100
gotify/server#1068 · 1 reaction ·
Maintainers usually reply within 2 days