[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecar
Maintainers usually reply within 1 day
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 78/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- go, kubernetes
- Domain
- infrastructure
Research direction
Start with gpuPodSpecFilter in cmd/gpu-operator/main.go, which the issue identifies as iterating only pod.Spec.Containers. Check how it identifies GPU resource requests and whether native sidecars in pod.Spec.InitContainers are covered. Done when a pod declaring nvidia.com/gpu only in a native sidecar is selected for deletion during driver upgrade; the issue does not name a test file.
Written by the indexing model from the issue text.
Description
Describe the bug
During an automatic driver upgrade, the upgrade controller does not delete workload pods whose nvidia.com/gpu resource is declared only in a native sidecar container (an entry in spec.initContainers with restartPolicy: Always).
The GPU pod filter, gpuPodSpecFilter in cmd/gpu-operator/main.go, only iterates pod.Spec.Containers, so such pods are treated as non-GPU pods ("No pods require deletion" logged). They keep the GPU open, k8s-driver-manager cannot unload the driver, the driver pod goes into Init:CrashLoopBackOff and the node ends up in upgrade-failed.
To Reproduce
- Run a Deployment whose main container has no GPU request, and a native sidecar with one:
initContainers:
- name: gpu-sidecar
restartPolicy: Always
resources:
limits:
nvidia.com/gpu: 1
- Change driver.version in the ClusterPolicy.
- Observe the operator log and the driver pod.
Expected behavior
Pods that declare nvidia.com/gpu* in resources of any container, including native sidecars, are selected by gpuPodSpecFilter and deleted in pod-deletion-required.
Environment (please provide the following information):
- GPU Operator Version: v26.7.1
- OS: Ubuntu24.04
- Kernel Version: 6.17.0-19-generic
- Container Runtime Version: containerd 2.2.6
- Kubernetes Distro and Version: K8s v1.36.1
- Dominant language
- Go
- Stars
- 2.9k
- Forks
- 569
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 76
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/gpu-operator
-
[Bug]: GPUCluster common name label breaks DRA validator selectorPossibly taken @ajavanma claimed this 18 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2955 · 1 comment ·
Maintainers usually reply within 1 day
-
feature lifecycle/frozen needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 25/100
NVIDIA/gpu-operator#3035 · 1 reaction ·
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/gpu-operator#3027 ·
Maintainers usually reply within 1 day
-
Ensure automated backport commits have verified signaturesPossibly taken @asivanadi0 claimed this 9 days ago. Opengood-first-issue
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/gpu-operator#2997 · 1 comment ·
Maintainers usually reply within 1 day
-
[Bug]: Latest Nvidia GPU Operator v26.7.1 reports large numbers of critical and high CVEs in Trivy scan outputPossibly taken @rahulait claimed this 10 days ago. Openmore-information-needed needs-triage
NVIDIA/gpu-operator#2991 · 3 comments · 1 assignee ·
Maintainers usually reply within 1 day
All issues in NVIDIA/gpu-operator
Similar issues
-
bug frontend good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
oalders/clodhopper#133 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
peasant-labs/peasant#596 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
hatchet-dev/hatchet#5179 ·
Maintainers usually reply within 1 day
-
bug triage
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
FairwindsOps/nova#484 ·