[Bug]: RHOCP CRI-O SELinux: GSP firmware read denial in workload context leads to RmInitAdapter failure
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- go, kubernetes
- Ambito
- infrastructure, security
Direzione di ricerca
Start by tracing how the driver container creates and labels /run/nvidia/driver/lib/firmware; the report does not identify the responsible component or file. Compare that path with the existing container_file_t labeling behavior and check the OpenShift driver-container configuration. Reproduce the late-GPU attach and verify the firmware label and workload AVC; done means a late-attached GPU initializes without a custom broad SELinux policy.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
On OpenShift with SELinux enforcing, a GPU that appears on the node after the driver container has finished installing cannot be initialized for compute. The kernel fails to load gsp_tu10x.bin with error -13 (EACCES). The cause is an SELinux denial: the workload process (container_t) is not allowed to read the firmware file, which is labeled container_var_run_t.
The rest of the driver rootfs under /run/nvidia/driver is labeled container_file_t, which workloads can read. Only the firmware directory has a different label.
GPUs that were present when the driver container installed work. They are initialized by privileged components (spc_t), so the label does not matter to them. A GPU that appears later is first touched by the workload, which is denied.
Our GPUs are attached dynamically by CoHDI (composable GPUs over a PCIe fabric). This report is not about CoHDI itself. It just creates the situation where a GPU first appears after the driver install.
Visible symptoms: the GPU shows up in nvidia-smi in the driver pod with 0 MiB used and no processes. Node Feature Discovery and GFD still see and label it. The workload does not crash on its own.
To Reproduce
- OpenShift 4.22.9 worker, SELinux enforcing, GPU Operator v26.3.3, no GPU attached yet.
- Attach GPU1 (we create a pod that requests 1 GPU through CoHDI). The driver container installs and loads the modules. The workload uses GPU1 without problems.
- While pod1 is running, attach GPU2 (we create a second pod that requests a GPU).
- The second pod starts, but GPU2 is not usable. The kernel log shows the firmware load failing, and
auditctl/ausearchshows the AVC denial forpythonincontainer_t. - Detaching and reattaching GPU1 later fails the same way.
We have not tried this without CoHDI. Any way of making a GPU appear after the driver container started should show the same thing, but that is untested.
Expected behavior
A GPU that appears after the driver container started can be initialized and used by a normal workload, with no extra SELinux policy. For example, the firmware files carry a label that container_t can read, or a privileged component initializes the new GPU first.
Environment
- GPU Operator Version: v26.3.3
- NVIDIA DRA driver for GPUs: v0.4.1
- NVIDIA driver: 580.126.20 (open kernel module), GPU: NVIDIA A30
- OS: RHCOS (OpenShift 4.22.9)
- Kernel Version: 5.14.0-687.35.1.el9_8.x86_64
- Container Runtime Version: CRI-O 1.35.5
- Kubernetes Distro and Version: OpenShift (RHOCP) 4.22.9
Evidence
The SELinux workaround module was removed before the run.
1. The workload is denied on the firmware file
type=SYSCALL msg=audit(...): syscall=257 success=no exit=-5
comm="python" exe="/usr/bin/python3.12"
subj=system_u:system_r:container_t:s0:c7,c28
type=AVC msg=audit(...): avc: denied { read }
comm="python" name="gsp_tu10x.bin" dev="tmpfs" ino=14990
scontext=system_u:system_r:container_t:s0:c7,c28
tcontext=system_u:object_r:container_var_run_t:s0 tclass=file permissive=0
The same denial repeats a few milliseconds later. The GPU1 attach in the same run produced zero AVC denials. The device node itself is readable: /dev/nvidia1 carries container_file_t:s0:c7,c28.
2. Kernel log, same second as the denial
kernel: nvidia 0000:96:00.0: loading /run/nvidia/driver/lib/firmware/nvidia/580.126.20/gsp_tu10x.bin failed with error -13
kernel: nvidia 0000:96:00.0: Direct firmware load for nvidia/580.126.20/gsp_tu10x.bin failed with error -2
kernel: NVRM: RmFetchGspRmImages: No firmware image found
kernel: NVRM: GPU 0000:96:00.0: RmInitAdapter failed! (0x61:0x56:1914)
kernel: NVRM: GPU 0000:96:00.0: rm_init_adapter failed, device minor number 1
-13 is EACCES. -2 is the fallback path (ENOENT). There is no equivalent failure for GPU1 (0000:95:00.0).
3. The privileged path succeeds on the same files
Captured with auditctl -w on the firmware directory (logs allowed and denied accesses):
11:53:26 comm="nvidia-installe" subj=system_u:system_r:spc_t:s0 success=yes (creates gsp_tu10x.bin, gsp_ga10x.bin)
11:53:31 comm="nvidia-persiste" subj=system_u:system_r:spc_t:s0 success=yes (opens gsp_tu10x.bin)
The same pair, in the same order, was seen in two separate runs.
4. It is not tied to "GPU2". A reattached GPU1 fails the same way
In an earlier test (2026-09-15), GPU1 first succeeded (22379 MiB used, active python process, no denial). After it was detached and reattached, the same denial appeared, with the same file and label:
Tue Sep 15 10:26:18 2026 avc: denied { read } for pid=341212 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
Tue Sep 15 10:50:14 2026 avc: denied { read } for pid=380168 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
So the variable is "was the device present when the driver container installed", not the slot or the order.
5. Label mismatch inside the same mount
$ ls -Z /run/nvidia/driver/lib/firmware/nvidia/580.126.20/
system_u:object_r:container_var_run_t:s0 gsp_ga10x.bin
system_u:object_r:container_var_run_t:s0 gsp_tu10x.bin
$ ls -Z /run/nvidia/driver/lib64/ | head
system_u:object_r:container_file_t:s0 LLVMgold.so
system_u:object_r:container_file_t:s0 Mcrt1.o
...
container_t can read container_file_t under the base policy. It cannot read container_var_run_t. Same result in all three runs. The mount is on tmpfs and does not exist before the driver container installs.
6. Driver modules load once per node
The nvidia, nvidia-modeset and nvidia-nvlink modules load once, after GPU1 appears. No module load happens when GPU2 appears six minutes later. So the GSP firmware load is per device, on first use.
7. Kernel attach events (context only)
pciehp: Slot(1): Card present / Link Up -> 0000:95:00.0 [10de:20b7]
pciehp: Slot(2): Card present / Link Up -> 0000:96:00.0 [10de:20b7]
Other platforms we tested
Same test on each (attach GPU1, then attach GPU2 while the GPU1 workload runs). FACT = seen in the logs we collected. HYP = our assumption. UNKNOWN = the capture failed or was not taken.
| Area | Ubuntu + containerd | RHEL + containerd | RKE2 (SLES 16) | OpenShift (CRI-O) |
|---|---|---|---|---|
| GPU2 usable after late attach | FACT: yes (79 GB used) | FACT: yes | FACT: yes | FACT: no, RmInitAdapter failed |
| GPU model / driver | A100 80GB / 580.126.20 | A100 80GB / 580.126.20 | A100 80GB / 595.71.05 | A30 / 580.126.20 |
| SELinux | UNKNOWN: not captured | FACT: enforcing | FACT: enabled (driver log). Mode not captured | FACT: enforcing |
| Runtime SELinux support | UNKNOWN | FACT: containerd enableSelinux: true |
UNKNOWN | FACT: enabled |
Denial on gsp_*.bin |
UNKNOWN: no audit data | UNKNOWN: audit watch failed, pipeline audit files empty. A short host AVC sample has no firmware denial | FACT: ausearch returned no matches. UNKNOWN whether it covered the attach |
FACT: repeated, reproducible |
Kernel -13 / RmInitAdapter |
UNKNOWN: no kernel log | FACT: none in the host journal we have. It does not cover the test window | FACT: none in the kernel log | FACT: present, same second as the denial |
| Label of the firmware directory | UNKNOWN | UNKNOWN: ls -Z returned "Permission denied" |
UNKNOWN: script used a hardcoded 580.126.20 path, node has 595.71.05 | FACT: container_var_run_t |
Privileged nvidia-smi before the workload |
Not done | Not done | Skipped (NODE not set) |
Not done |
| Who reads the firmware first on GPU2 | UNKNOWN | UNKNOWN | UNKNOWN | FACT: the workload (python, container_t) |
What we take from this, and what we do not:
- GPU2 works on the three other platforms. We do not know why. It could be a different firmware label, or a different first reader.
- The driver container script calls
chcon -R -t container_file_t /run/nvidia/driver/devwhen SELinux is enabled (seen in the RKE2 log, and the same step is printed in the RHEL log). It does not relabel the firmware directory. We have not checked whether the OpenShift driver container uses the same script. - The driver message
knvlinkCoreShutdownDeviceLinks_IMPL ... for GPU1also appears on RKE2 and RHEL runs where GPU2 works. We think it is unrelated. - We are re-running RHEL and RKE2 with working captures and will update this issue.
Confirmed vs. not confirmed
Confirmed by logs on this OpenShift setup:
- The firmware file is labeled
container_var_run_t, the workload is denied, and the kernel firmware load fails with-13at the same second. spc_tcomponents (installer, persistenced) read the same file without problems.- The behavior reproduced in three runs, and a reattached GPU1 fails the same way as GPU2.
- With a local SELinux module that allows
container_tto readcontainer_var_run_t, the same steps show zero denials.
Not confirmed yet:
- The firmware label and first reader on other platforms (see the table).
- Whether a privileged
nvidia-smirun before the workload would initialize a late-attached GPU and avoid the problem. Not tested yet. It would confirm that "first process to open the device" is the trigger. - Whether
nvidia-persistencedpicks up devices that show up after its start. - Which component creates
/run/nvidia/driver/lib/firmwareand sets its label inside the driver container.
Workaround (works, not a fix)
A custom policy module allowing container_t to read container_var_run_t files. In a run where it was left installed by mistake, the same attach steps produced no denials. Illustrative form (our local module may differ in detail):
allow container_t container_var_run_t:file { read open getattr };
This widens access for all container_t workloads, which is why we would like a proper fix.
Possible fixes
- Label: make the firmware directory use
container_file_t, like the rest of/run/nvidia/driver. - Initialization: make sure a privileged component initializes GPUs that appear after the driver install, so the first firmware read never comes from an unprivileged workload.
Information to attach
- kubernetes pods status:
kubectl get pods -n gpu-operator(attach) - kubernetes daemonset status:
kubectl get ds -n gpu-operator(attach) - pod/ds in error state: not applicable, all pods are Running. The failure is inside the workload's GPU init.
-
nvidia-smifrom the driver container:kubectl exec DRIVER_POD_NAME -n gpu-operator -c nvidia-driver-ctr -- nvidia-smi(attach, shows GPU2 with 0 MiB and no processes) - Audit log from the
auditctl -wwatch (keygsp_access_watch) - Full kernel log for the run (
journalctl -k) -
ls -Zoutput of the firmware directory andlib64 - Test scripts used for the attach and reattach steps
- must-gather: can be sent to [email protected] if you want it
- Lingua principale
- Go
- Stelle
- 2.9k
- Fork
- 569
- Merge medio
- 1g 10h
- PR unite (30g)
- 78
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/gpu-operator
-
[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecarForse già presa Una pull request collegata a questa issue è aperta o già unita. Apertabug needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
NVIDIA/gpu-operator#3026 ·
I maintainer di solito rispondono entro 1 giorno
-
Make NVIDIADriver node-pool rendering deterministicForse già presa @efegokdemir l’ha presa 9 giorni fa. Apertadsx-ws-0930 good-first-issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
NVIDIA/gpu-operator#2981 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: GPUCluster common name label breaks DRA validator selectorForse già presa @ajavanma l’ha presa 17 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
NVIDIA/gpu-operator#2955 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Ensure automated backport commits have verified signaturesForse già presa @asivanadi0 l’ha presa 8 giorni fa. Apertagood-first-issue
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
NVIDIA/gpu-operator#2997 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: Latest Nvidia GPU Operator v26.7.1 reports large numbers of critical and high CVEs in Trivy scan outputForse già presa @rahulait l’ha presa 9 giorni fa. Apertamore-information-needed needs-triage
NVIDIA/gpu-operator#2991 · 3 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di NVIDIA/gpu-operator
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
mvanhorn/cli-printing-press#4980 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
priority: p3 type: feat
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
googleapis/librarian#7775 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
OpenTollGate/tollgate-module-basic-go#833 ·
I maintainer di solito rispondono entro 1 giorno
-
documentation
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
I maintainer di solito rispondono entro 1 giorno