Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecar

Open Beginner friendly
#3,026 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@shin4141 is already working on this.

Since Oct 8, 2026.

  • #3028 by @shin4141 — open

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
go, kubernetes

Research direction

Start with gpuPodSpecFilter in cmd/gpu-operator/main.go, which the issue identifies as iterating only pod.Spec.Containers. Check how it identifies GPU resource requests and whether native sidecars in pod.Spec.InitContainers are covered. Done when a pod declaring nvidia.com/gpu only in a native sidecar is selected for deletion during driver upgrade; the issue does not name a test file.

Written by the indexing model from the issue text.

Description

bug needs-triage

Describe the bug
During an automatic driver upgrade, the upgrade controller does not delete workload pods whose nvidia.com/gpu resource is declared only in a native sidecar container (an entry in spec.initContainers with restartPolicy: Always).
The GPU pod filter, gpuPodSpecFilter in cmd/gpu-operator/main.go, only iterates pod.Spec.Containers, so such pods are treated as non-GPU pods ("No pods require deletion" logged). They keep the GPU open, k8s-driver-manager cannot unload the driver, the driver pod goes into Init:CrashLoopBackOff and the node ends up in upgrade-failed.

To Reproduce

  1. Run a Deployment whose main container has no GPU request, and a native sidecar with one:
  initContainers:
  - name: gpu-sidecar
    restartPolicy: Always
    resources:
      limits: 
        nvidia.com/gpu: 1
  1. Change driver.version in the ClusterPolicy.
  2. Observe the operator log and the driver pod.

Expected behavior
Pods that declare nvidia.com/gpu* in resources of any container, including native sidecars, are selected by gpuPodSpecFilter and deleted in pod-deletion-required.

Environment (please provide the following information):

  • GPU Operator Version: v26.7.1
  • OS: Ubuntu24.04
  • Kernel Version: 6.17.0-19-generic
  • Container Runtime Version: containerd 2.2.6
  • Kubernetes Distro and Version: K8s v1.36.1
Dominant language
Go
Stars
2.9k
Forks
569
Avg merge
1d 7h
Merged PRs (30d)
76

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/gpu-operator

All issues in NVIDIA/gpu-operator

Similar issues

More Go issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.