Fix ComputeDomain daemon permissions on OpenShift
Maintainers usually reply within 1 day
Assessment
This issue has not been assessed yet.
Description
Description:
When testing MNNVL on OpenShift 4.22.1 with NVL72 nodes and GPU Operator v26.7.1, ComputeDomain daemon Pods encounter two permission failures:
- Incorrect SCC selection: Pods are admitted under restricted-v2 and crash because they cannot write /imexd/imexd.cfg. Although the daemon’s service account is listed in the nvidia-dra-driver SCC, that SCC has no priority, allowing OpenShift to prefer restricted-v2.
- Missing RBAC permission: The daemon’s ClusterRole lacks delete on computedomaincliques.resource.nvidia.com. Owner-reference admission rejects its requests with:
cannot set an ownerRef on a resource you can't delete
Expected behavior:
ComputeDomain daemon Pods should start successfully and manage clique owner references.
Proposed fix:
Give the custom DRA SCC a positive priority and add the missing clique permission to both the daemon’s ClusterRole and the Operator’s Helm/OLM RBAC.
- Dominant language
- Go
- Stars
- 2.9k
- Forks
- 569
- Avg merge
- 1d 10h
- Merged PRs (30d)
- 78
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/gpu-operator
-
[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecarPossibly taken A pull request linked to this issue is open or already merged. Openbug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/gpu-operator#3026 ·
Maintainers usually reply within 1 day
-
Make NVIDIADriver node-pool rendering deterministicPossibly taken @efegokdemir claimed this 9 days ago. Opendsx-ws-0930 good-first-issue
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2981 · 1 comment ·
Maintainers usually reply within 1 day
-
[Bug]: GPUCluster common name label breaks DRA validator selectorPossibly taken @ajavanma claimed this 16 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2955 · 1 comment ·
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/gpu-operator#3027 ·
Maintainers usually reply within 1 day
-
Ensure automated backport commits have verified signaturesPossibly taken @asivanadi0 claimed this 8 days ago. Opengood-first-issue
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/gpu-operator#2997 · 1 comment ·
Maintainers usually reply within 1 day
All issues in NVIDIA/gpu-operator
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
mvanhorn/cli-printing-press#4980 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
priority: p3 type: feat
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
googleapis/librarian#7775 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
OpenTollGate/tollgate-module-basic-go#833 ·
Maintainers usually reply within 1 day
-
documentation
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day