[Bug]: RHOCP CRI-O SELinux: GSP firmware read denial in workload context leads to RmInitAdapter failure
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- go, kubernetes
- Área
- infrastructure, security
Línea de trabajo
Start by tracing how the driver container creates and labels /run/nvidia/driver/lib/firmware; the report does not identify the responsible component or file. Compare that path with the existing container_file_t labeling behavior and check the OpenShift driver-container configuration. Reproduce the late-GPU attach and verify the firmware label and workload AVC; done means a late-attached GPU initializes without a custom broad SELinux policy.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
On OpenShift with SELinux enforcing, a GPU that appears on the node after the driver container has finished installing cannot be initialized for compute. The kernel fails to load gsp_tu10x.bin with error -13 (EACCES). The cause is an SELinux denial: the workload process (container_t) is not allowed to read the firmware file, which is labeled container_var_run_t.
The rest of the driver rootfs under /run/nvidia/driver is labeled container_file_t, which workloads can read. Only the firmware directory has a different label.
GPUs that were present when the driver container installed work. They are initialized by privileged components (spc_t), so the label does not matter to them. A GPU that appears later is first touched by the workload, which is denied.
Our GPUs are attached dynamically by CoHDI (composable GPUs over a PCIe fabric). This report is not about CoHDI itself. It just creates the situation where a GPU first appears after the driver install.
Visible symptoms: the GPU shows up in nvidia-smi in the driver pod with 0 MiB used and no processes. Node Feature Discovery and GFD still see and label it. The workload does not crash on its own.
To Reproduce
- OpenShift 4.22.9 worker, SELinux enforcing, GPU Operator v26.3.3, no GPU attached yet.
- Attach GPU1 (we create a pod that requests 1 GPU through CoHDI). The driver container installs and loads the modules. The workload uses GPU1 without problems.
- While pod1 is running, attach GPU2 (we create a second pod that requests a GPU).
- The second pod starts, but GPU2 is not usable. The kernel log shows the firmware load failing, and
auditctl/ausearchshows the AVC denial forpythonincontainer_t. - Detaching and reattaching GPU1 later fails the same way.
We have not tried this without CoHDI. Any way of making a GPU appear after the driver container started should show the same thing, but that is untested.
Expected behavior
A GPU that appears after the driver container started can be initialized and used by a normal workload, with no extra SELinux policy. For example, the firmware files carry a label that container_t can read, or a privileged component initializes the new GPU first.
Environment
- GPU Operator Version: v26.3.3
- NVIDIA DRA driver for GPUs: v0.4.1
- NVIDIA driver: 580.126.20 (open kernel module), GPU: NVIDIA A30
- OS: RHCOS (OpenShift 4.22.9)
- Kernel Version: 5.14.0-687.35.1.el9_8.x86_64
- Container Runtime Version: CRI-O 1.35.5
- Kubernetes Distro and Version: OpenShift (RHOCP) 4.22.9
Evidence
The SELinux workaround module was removed before the run.
1. The workload is denied on the firmware file
type=SYSCALL msg=audit(...): syscall=257 success=no exit=-5
comm="python" exe="/usr/bin/python3.12"
subj=system_u:system_r:container_t:s0:c7,c28
type=AVC msg=audit(...): avc: denied { read }
comm="python" name="gsp_tu10x.bin" dev="tmpfs" ino=14990
scontext=system_u:system_r:container_t:s0:c7,c28
tcontext=system_u:object_r:container_var_run_t:s0 tclass=file permissive=0
The same denial repeats a few milliseconds later. The GPU1 attach in the same run produced zero AVC denials. The device node itself is readable: /dev/nvidia1 carries container_file_t:s0:c7,c28.
2. Kernel log, same second as the denial
kernel: nvidia 0000:96:00.0: loading /run/nvidia/driver/lib/firmware/nvidia/580.126.20/gsp_tu10x.bin failed with error -13
kernel: nvidia 0000:96:00.0: Direct firmware load for nvidia/580.126.20/gsp_tu10x.bin failed with error -2
kernel: NVRM: RmFetchGspRmImages: No firmware image found
kernel: NVRM: GPU 0000:96:00.0: RmInitAdapter failed! (0x61:0x56:1914)
kernel: NVRM: GPU 0000:96:00.0: rm_init_adapter failed, device minor number 1
-13 is EACCES. -2 is the fallback path (ENOENT). There is no equivalent failure for GPU1 (0000:95:00.0).
3. The privileged path succeeds on the same files
Captured with auditctl -w on the firmware directory (logs allowed and denied accesses):
11:53:26 comm="nvidia-installe" subj=system_u:system_r:spc_t:s0 success=yes (creates gsp_tu10x.bin, gsp_ga10x.bin)
11:53:31 comm="nvidia-persiste" subj=system_u:system_r:spc_t:s0 success=yes (opens gsp_tu10x.bin)
The same pair, in the same order, was seen in two separate runs.
4. It is not tied to "GPU2". A reattached GPU1 fails the same way
In an earlier test (2026-09-15), GPU1 first succeeded (22379 MiB used, active python process, no denial). After it was detached and reattached, the same denial appeared, with the same file and label:
Tue Sep 15 10:26:18 2026 avc: denied { read } for pid=341212 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
Tue Sep 15 10:50:14 2026 avc: denied { read } for pid=380168 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
So the variable is "was the device present when the driver container installed", not the slot or the order.
5. Label mismatch inside the same mount
$ ls -Z /run/nvidia/driver/lib/firmware/nvidia/580.126.20/
system_u:object_r:container_var_run_t:s0 gsp_ga10x.bin
system_u:object_r:container_var_run_t:s0 gsp_tu10x.bin
$ ls -Z /run/nvidia/driver/lib64/ | head
system_u:object_r:container_file_t:s0 LLVMgold.so
system_u:object_r:container_file_t:s0 Mcrt1.o
...
container_t can read container_file_t under the base policy. It cannot read container_var_run_t. Same result in all three runs. The mount is on tmpfs and does not exist before the driver container installs.
6. Driver modules load once per node
The nvidia, nvidia-modeset and nvidia-nvlink modules load once, after GPU1 appears. No module load happens when GPU2 appears six minutes later. So the GSP firmware load is per device, on first use.
7. Kernel attach events (context only)
pciehp: Slot(1): Card present / Link Up -> 0000:95:00.0 [10de:20b7]
pciehp: Slot(2): Card present / Link Up -> 0000:96:00.0 [10de:20b7]
Other platforms we tested
Same test on each (attach GPU1, then attach GPU2 while the GPU1 workload runs). FACT = seen in the logs we collected. HYP = our assumption. UNKNOWN = the capture failed or was not taken.
| Area | Ubuntu + containerd | RHEL + containerd | RKE2 (SLES 16) | OpenShift (CRI-O) |
|---|---|---|---|---|
| GPU2 usable after late attach | FACT: yes (79 GB used) | FACT: yes | FACT: yes | FACT: no, RmInitAdapter failed |
| GPU model / driver | A100 80GB / 580.126.20 | A100 80GB / 580.126.20 | A100 80GB / 595.71.05 | A30 / 580.126.20 |
| SELinux | UNKNOWN: not captured | FACT: enforcing | FACT: enabled (driver log). Mode not captured | FACT: enforcing |
| Runtime SELinux support | UNKNOWN | FACT: containerd enableSelinux: true |
UNKNOWN | FACT: enabled |
Denial on gsp_*.bin |
UNKNOWN: no audit data | UNKNOWN: audit watch failed, pipeline audit files empty. A short host AVC sample has no firmware denial | FACT: ausearch returned no matches. UNKNOWN whether it covered the attach |
FACT: repeated, reproducible |
Kernel -13 / RmInitAdapter |
UNKNOWN: no kernel log | FACT: none in the host journal we have. It does not cover the test window | FACT: none in the kernel log | FACT: present, same second as the denial |
| Label of the firmware directory | UNKNOWN | UNKNOWN: ls -Z returned "Permission denied" |
UNKNOWN: script used a hardcoded 580.126.20 path, node has 595.71.05 | FACT: container_var_run_t |
Privileged nvidia-smi before the workload |
Not done | Not done | Skipped (NODE not set) |
Not done |
| Who reads the firmware first on GPU2 | UNKNOWN | UNKNOWN | UNKNOWN | FACT: the workload (python, container_t) |
What we take from this, and what we do not:
- GPU2 works on the three other platforms. We do not know why. It could be a different firmware label, or a different first reader.
- The driver container script calls
chcon -R -t container_file_t /run/nvidia/driver/devwhen SELinux is enabled (seen in the RKE2 log, and the same step is printed in the RHEL log). It does not relabel the firmware directory. We have not checked whether the OpenShift driver container uses the same script. - The driver message
knvlinkCoreShutdownDeviceLinks_IMPL ... for GPU1also appears on RKE2 and RHEL runs where GPU2 works. We think it is unrelated. - We are re-running RHEL and RKE2 with working captures and will update this issue.
Confirmed vs. not confirmed
Confirmed by logs on this OpenShift setup:
- The firmware file is labeled
container_var_run_t, the workload is denied, and the kernel firmware load fails with-13at the same second. spc_tcomponents (installer, persistenced) read the same file without problems.- The behavior reproduced in three runs, and a reattached GPU1 fails the same way as GPU2.
- With a local SELinux module that allows
container_tto readcontainer_var_run_t, the same steps show zero denials.
Not confirmed yet:
- The firmware label and first reader on other platforms (see the table).
- Whether a privileged
nvidia-smirun before the workload would initialize a late-attached GPU and avoid the problem. Not tested yet. It would confirm that "first process to open the device" is the trigger. - Whether
nvidia-persistencedpicks up devices that show up after its start. - Which component creates
/run/nvidia/driver/lib/firmwareand sets its label inside the driver container.
Workaround (works, not a fix)
A custom policy module allowing container_t to read container_var_run_t files. In a run where it was left installed by mistake, the same attach steps produced no denials. Illustrative form (our local module may differ in detail):
allow container_t container_var_run_t:file { read open getattr };
This widens access for all container_t workloads, which is why we would like a proper fix.
Possible fixes
- Label: make the firmware directory use
container_file_t, like the rest of/run/nvidia/driver. - Initialization: make sure a privileged component initializes GPUs that appear after the driver install, so the first firmware read never comes from an unprivileged workload.
Information to attach
- kubernetes pods status:
kubectl get pods -n gpu-operator(attach) - kubernetes daemonset status:
kubectl get ds -n gpu-operator(attach) - pod/ds in error state: not applicable, all pods are Running. The failure is inside the workload's GPU init.
-
nvidia-smifrom the driver container:kubectl exec DRIVER_POD_NAME -n gpu-operator -c nvidia-driver-ctr -- nvidia-smi(attach, shows GPU2 with 0 MiB and no processes) - Audit log from the
auditctl -wwatch (keygsp_access_watch) - Full kernel log for the run (
journalctl -k) -
ls -Zoutput of the firmware directory andlib64 - Test scripts used for the attach and reattach steps
- must-gather: can be sent to [email protected] if you want it
- Lenguaje dominante
- Go
- Estrellas
- 2.9k
- Forks
- 569
- Merge medio
- 1 d 10 h
- PR fusionados (30 d)
- 78
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA/gpu-operator
-
[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecarPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abiertobug needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
NVIDIA/gpu-operator#3026 ·
Los mantenedores suelen responder en 1 día
-
Make NVIDIADriver node-pool rendering deterministicPosiblemente ocupada @efegokdemir la tomó hace 9 días. Abiertodsx-ws-0930 good-first-issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
NVIDIA/gpu-operator#2981 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[Bug]: GPUCluster common name label breaks DRA validator selectorPosiblemente ocupada @ajavanma la tomó hace 16 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
NVIDIA/gpu-operator#2955 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Ensure automated backport commits have verified signaturesPosiblemente ocupada @asivanadi0 la tomó hace 7 días. Abiertogood-first-issue
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
NVIDIA/gpu-operator#2997 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[Bug]: Latest Nvidia GPU Operator v26.7.1 reports large numbers of critical and high CVEs in Trivy scan outputPosiblemente ocupada @rahulait la tomó hace 8 días. Abiertomore-information-needed needs-triage
NVIDIA/gpu-operator#2991 · 3 comentarios · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de NVIDIA/gpu-operator
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
siyuan-note/siyuan#20353 ·
Los mantenedores suelen responder en 1 día
-
attributes-natural-language "en-US" is rejected by PAPPL >= 1.4.12 printers (RFC 8011 requires lowercase)Posiblemente ocupada @ChrisEdgington la tomó hoy. Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 84/100
OpenPrinting/ipp-usb#140 ·
-
Discriminator mapping keys are listed in a random orderPosiblemente ocupada @reuvenharrison la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Idle compaction monitors LIST the replica every tick when the newest destination file spans more than one TXIDPosiblemente ocupada @pishuv la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
benbjohnson/litestream#1563 ·
Los mantenedores suelen responder en 2 días
-
triage needed
Dificultad 1/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 2 días