[Bug]: GPUCluster ignores daemonsets.updateStrategy=OnDelete
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- go, kubernetes
- Domain
- infrastructure
Research direction
Start with the four component getManifestObjects helpers for the DRA kubelet plugin, DRA validator, DCGM, and DCGM exporter, then read controllers/object_controls.go for the existing ClusterPolicy OnDelete revision checks. Verify rendered DaemonSets for omitted, RollingUpdate, and OnDelete strategies, including leftover rolling-update settings. Done means OnDelete omits rollingUpdate fields, RollingUpdate preserves current output, and readiness remains correct across the stated pod-replacement cases.
Written by the indexing model from the issue text.
Description
Describe the bug
GPUCluster.spec.daemonsets.updateStrategy: OnDelete is accepted by the CRD, but the generated DaemonSets use RollingUpdate for the DRA kubelet plugin, DRA validator, DCGM and DCGM exporter. Consequently, the rendered resources allow automatic replacement on a pod-template change despite the requested manual update strategy.
The four templates hard-code RollingUpdate. Their render data already includes Daemonsets.UpdateStrategy, and the GPUCluster apply path does not run the ClusterPolicy strategy transformer.
To Reproduce
At main revision 4fdfb1db7ddb87ba8969c52c634a5aeff104ab5f, use this GPUCluster configuration fragment (enable DCGM to include that optional component):
spec:
daemonsets:
updateStrategy: OnDelete
dcgm:
enabled: true
Render the four component manifests through their existing getManifestObjects helpers. Each DaemonSet has spec.updateStrategy.type: RollingUpdate, although the common configuration requests OnDelete. The DRA plugin and validator also include rollingUpdate.maxUnavailable: "100%".
Expected behavior
All four DaemonSets should honor OnDelete and omit rollingUpdate settings. An omitted or explicit RollingUpdate strategy should preserve the current output, including the DRA components' 100% maximum unavailable setting.
The proposed scope keeps readiness semantics unchanged: healthy old pods do not satisfy the requested configuration, so GPUCluster stays NotReady until an administrator replaces them and the updated pods are available. This follows the existing ClusterPolicy OnDelete revision checks in controllers/object_controls.go. The GPUCluster state manager continues reconciling every component while a preceding one is NotReady, so the other DaemonSets still receive their desired configuration. Maintainer feedback on this contract is welcome.
Environment (please provide the following information):
- GPU Operator: main at
4fdfb1db7ddb87ba8969c52c634a5aeff104ab5f. - OS/architecture: Linux ARM64; Go 1.27.1.
- Reproduction: real template renderer and typed Kubernetes objects; GPU runtime is not needed for this reproduction.
- Live GPU Operator deployment: not tested. Separate CPU Kubernetes lifecycle validation is recorded with the proposed PR.
Information to attach
The regression covers all four operands with omitted, explicit RollingUpdate, OnDelete, and OnDelete plus leftover rolling-update settings. Readiness cases cover old, partially replaced, updated-but-unavailable and fully updated pods under both strategies. Test results will accompany the proposed PR.
- Dominant language
- Go
- Stars
- 2.9k
- Forks
- 552
- Avg merge
- 1d 21h
- Merged PRs (30d)
- 78
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/gpu-operator
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NVIDIA/gpu-operator#2968 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2955 · 1 comment ·
Maintainers usually reply within 1 day
-
lifecycle/stale question
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
NVIDIA/gpu-operator#2280 · 2 comments · 1 reaction ·
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 3/5 1-2 days Newbie friendliness 68/100
NVIDIA/gpu-operator#2970 · 2 comments ·
Maintainers usually reply within 1 day
-
feature lifecycle/frozen needs-triage
Difficulty 3/5 1-2 days Newbie friendliness 65/100
NVIDIA/gpu-operator#2938 · 1 comment ·
Maintainers usually reply within 1 day
All issues in NVIDIA/gpu-operator
Similar issues
-
bug needs triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
netdata/netdata#24062 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
meshery/meshery#22119 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
automation documentation
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day