Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Bug]: RHOCP CRI-O SELinux: GSP firmware read denial in workload context leads to RmInitAdapter failure

Abierto
#3,027 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
48/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
go, kubernetes

Línea de trabajo

Start by tracing how the driver container creates and labels /run/nvidia/driver/lib/firmware; the report does not identify the responsible component or file. Compare that path with the existing container_file_t labeling behavior and check the OpenShift driver-container configuration. Reproduce the late-GPU attach and verify the firmware label and workload AVC; done means a late-attached GPU initializes without a custom broad SELinux policy.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

bug needs-triage

On OpenShift with SELinux enforcing, a GPU that appears on the node after the driver container has finished installing cannot be initialized for compute. The kernel fails to load gsp_tu10x.bin with error -13 (EACCES). The cause is an SELinux denial: the workload process (container_t) is not allowed to read the firmware file, which is labeled container_var_run_t.

The rest of the driver rootfs under /run/nvidia/driver is labeled container_file_t, which workloads can read. Only the firmware directory has a different label.

GPUs that were present when the driver container installed work. They are initialized by privileged components (spc_t), so the label does not matter to them. A GPU that appears later is first touched by the workload, which is denied.

Our GPUs are attached dynamically by CoHDI (composable GPUs over a PCIe fabric). This report is not about CoHDI itself. It just creates the situation where a GPU first appears after the driver install.

Visible symptoms: the GPU shows up in nvidia-smi in the driver pod with 0 MiB used and no processes. Node Feature Discovery and GFD still see and label it. The workload does not crash on its own.

To Reproduce

  1. OpenShift 4.22.9 worker, SELinux enforcing, GPU Operator v26.3.3, no GPU attached yet.
  2. Attach GPU1 (we create a pod that requests 1 GPU through CoHDI). The driver container installs and loads the modules. The workload uses GPU1 without problems.
  3. While pod1 is running, attach GPU2 (we create a second pod that requests a GPU).
  4. The second pod starts, but GPU2 is not usable. The kernel log shows the firmware load failing, and auditctl/ausearch shows the AVC denial for python in container_t.
  5. Detaching and reattaching GPU1 later fails the same way.

We have not tried this without CoHDI. Any way of making a GPU appear after the driver container started should show the same thing, but that is untested.

Expected behavior

A GPU that appears after the driver container started can be initialized and used by a normal workload, with no extra SELinux policy. For example, the firmware files carry a label that container_t can read, or a privileged component initializes the new GPU first.

Environment

  • GPU Operator Version: v26.3.3
  • NVIDIA DRA driver for GPUs: v0.4.1
  • NVIDIA driver: 580.126.20 (open kernel module), GPU: NVIDIA A30
  • OS: RHCOS (OpenShift 4.22.9)
  • Kernel Version: 5.14.0-687.35.1.el9_8.x86_64
  • Container Runtime Version: CRI-O 1.35.5
  • Kubernetes Distro and Version: OpenShift (RHOCP) 4.22.9

Evidence

The SELinux workaround module was removed before the run.

1. The workload is denied on the firmware file

type=SYSCALL msg=audit(...): syscall=257 success=no exit=-5
  comm="python" exe="/usr/bin/python3.12"
  subj=system_u:system_r:container_t:s0:c7,c28
type=AVC msg=audit(...): avc: denied { read }
  comm="python" name="gsp_tu10x.bin" dev="tmpfs" ino=14990
  scontext=system_u:system_r:container_t:s0:c7,c28
  tcontext=system_u:object_r:container_var_run_t:s0 tclass=file permissive=0

The same denial repeats a few milliseconds later. The GPU1 attach in the same run produced zero AVC denials. The device node itself is readable: /dev/nvidia1 carries container_file_t:s0:c7,c28.

2. Kernel log, same second as the denial

kernel: nvidia 0000:96:00.0: loading /run/nvidia/driver/lib/firmware/nvidia/580.126.20/gsp_tu10x.bin failed with error -13
kernel: nvidia 0000:96:00.0: Direct firmware load for nvidia/580.126.20/gsp_tu10x.bin failed with error -2
kernel: NVRM: RmFetchGspRmImages: No firmware image found
kernel: NVRM: GPU 0000:96:00.0: RmInitAdapter failed! (0x61:0x56:1914)
kernel: NVRM: GPU 0000:96:00.0: rm_init_adapter failed, device minor number 1

-13 is EACCES. -2 is the fallback path (ENOENT). There is no equivalent failure for GPU1 (0000:95:00.0).

3. The privileged path succeeds on the same files

Captured with auditctl -w on the firmware directory (logs allowed and denied accesses):

11:53:26  comm="nvidia-installe"  subj=system_u:system_r:spc_t:s0  success=yes  (creates gsp_tu10x.bin, gsp_ga10x.bin)
11:53:31  comm="nvidia-persiste"  subj=system_u:system_r:spc_t:s0  success=yes  (opens gsp_tu10x.bin)

The same pair, in the same order, was seen in two separate runs.

4. It is not tied to "GPU2". A reattached GPU1 fails the same way

In an earlier test (2026-09-15), GPU1 first succeeded (22379 MiB used, active python process, no denial). After it was detached and reattached, the same denial appeared, with the same file and label:

Tue Sep 15 10:26:18 2026  avc: denied { read } for pid=341212 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
Tue Sep 15 10:50:14 2026  avc: denied { read } for pid=380168 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0

So the variable is "was the device present when the driver container installed", not the slot or the order.

5. Label mismatch inside the same mount

$ ls -Z /run/nvidia/driver/lib/firmware/nvidia/580.126.20/
system_u:object_r:container_var_run_t:s0 gsp_ga10x.bin
system_u:object_r:container_var_run_t:s0 gsp_tu10x.bin

$ ls -Z /run/nvidia/driver/lib64/ | head
system_u:object_r:container_file_t:s0 LLVMgold.so
system_u:object_r:container_file_t:s0 Mcrt1.o
...

container_t can read container_file_t under the base policy. It cannot read container_var_run_t. Same result in all three runs. The mount is on tmpfs and does not exist before the driver container installs.

6. Driver modules load once per node

The nvidia, nvidia-modeset and nvidia-nvlink modules load once, after GPU1 appears. No module load happens when GPU2 appears six minutes later. So the GSP firmware load is per device, on first use.

7. Kernel attach events (context only)

pciehp: Slot(1): Card present / Link Up    -> 0000:95:00.0 [10de:20b7]
pciehp: Slot(2): Card present / Link Up    -> 0000:96:00.0 [10de:20b7]

Other platforms we tested

Same test on each (attach GPU1, then attach GPU2 while the GPU1 workload runs). FACT = seen in the logs we collected. HYP = our assumption. UNKNOWN = the capture failed or was not taken.

Area Ubuntu + containerd RHEL + containerd RKE2 (SLES 16) OpenShift (CRI-O)
GPU2 usable after late attach FACT: yes (79 GB used) FACT: yes FACT: yes FACT: no, RmInitAdapter failed
GPU model / driver A100 80GB / 580.126.20 A100 80GB / 580.126.20 A100 80GB / 595.71.05 A30 / 580.126.20
SELinux UNKNOWN: not captured FACT: enforcing FACT: enabled (driver log). Mode not captured FACT: enforcing
Runtime SELinux support UNKNOWN FACT: containerd enableSelinux: true UNKNOWN FACT: enabled
Denial on gsp_*.bin UNKNOWN: no audit data UNKNOWN: audit watch failed, pipeline audit files empty. A short host AVC sample has no firmware denial FACT: ausearch returned no matches. UNKNOWN whether it covered the attach FACT: repeated, reproducible
Kernel -13 / RmInitAdapter UNKNOWN: no kernel log FACT: none in the host journal we have. It does not cover the test window FACT: none in the kernel log FACT: present, same second as the denial
Label of the firmware directory UNKNOWN UNKNOWN: ls -Z returned "Permission denied" UNKNOWN: script used a hardcoded 580.126.20 path, node has 595.71.05 FACT: container_var_run_t
Privileged nvidia-smi before the workload Not done Not done Skipped (NODE not set) Not done
Who reads the firmware first on GPU2 UNKNOWN UNKNOWN UNKNOWN FACT: the workload (python, container_t)

What we take from this, and what we do not:

  • GPU2 works on the three other platforms. We do not know why. It could be a different firmware label, or a different first reader.
  • The driver container script calls chcon -R -t container_file_t /run/nvidia/driver/dev when SELinux is enabled (seen in the RKE2 log, and the same step is printed in the RHEL log). It does not relabel the firmware directory. We have not checked whether the OpenShift driver container uses the same script.
  • The driver message knvlinkCoreShutdownDeviceLinks_IMPL ... for GPU1 also appears on RKE2 and RHEL runs where GPU2 works. We think it is unrelated.
  • We are re-running RHEL and RKE2 with working captures and will update this issue.

Confirmed vs. not confirmed

Confirmed by logs on this OpenShift setup:

  • The firmware file is labeled container_var_run_t, the workload is denied, and the kernel firmware load fails with -13 at the same second.
  • spc_t components (installer, persistenced) read the same file without problems.
  • The behavior reproduced in three runs, and a reattached GPU1 fails the same way as GPU2.
  • With a local SELinux module that allows container_t to read container_var_run_t, the same steps show zero denials.

Not confirmed yet:

  • The firmware label and first reader on other platforms (see the table).
  • Whether a privileged nvidia-smi run before the workload would initialize a late-attached GPU and avoid the problem. Not tested yet. It would confirm that "first process to open the device" is the trigger.
  • Whether nvidia-persistenced picks up devices that show up after its start.
  • Which component creates /run/nvidia/driver/lib/firmware and sets its label inside the driver container.

Workaround (works, not a fix)

A custom policy module allowing container_t to read container_var_run_t files. In a run where it was left installed by mistake, the same attach steps produced no denials. Illustrative form (our local module may differ in detail):

allow container_t container_var_run_t:file { read open getattr };

This widens access for all container_t workloads, which is why we would like a proper fix.

Possible fixes

  1. Label: make the firmware directory use container_file_t, like the rest of /run/nvidia/driver.
  2. Initialization: make sure a privileged component initializes GPUs that appear after the driver install, so the first firmware read never comes from an unprivileged workload.

Information to attach

  • kubernetes pods status: kubectl get pods -n gpu-operator (attach)
  • kubernetes daemonset status: kubectl get ds -n gpu-operator (attach)
  • pod/ds in error state: not applicable, all pods are Running. The failure is inside the workload's GPU init.
  • nvidia-smi from the driver container: kubectl exec DRIVER_POD_NAME -n gpu-operator -c nvidia-driver-ctr -- nvidia-smi (attach, shows GPU2 with 0 MiB and no processes)
  • Audit log from the auditctl -w watch (key gsp_access_watch)
  • Full kernel log for the run (journalctl -k)
  • ls -Z output of the firmware directory and lib64
  • Test scripts used for the attach and reattach steps
  • must-gather: can be sent to [email protected] if you want it

rhocp-nvidia-issue-2026-10-08.zip

Lenguaje dominante
Go
Estrellas
2.9k
Forks
569
Merge medio
1 d 10 h
PR fusionados (30 d)
78

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de NVIDIA/gpu-operator

Todos los issues de NVIDIA/gpu-operator

Issues similares

Más issues de Go

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.