[Bug]: RHOCP CRI-O SELinux: GSP firmware read denial in workload context leads to RmInitAdapter failure
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- go, kubernetes
- Domain
- infrastructure, security
Research direction
Start by tracing how the driver container creates and labels /run/nvidia/driver/lib/firmware; the report does not identify the responsible component or file. Compare that path with the existing container_file_t labeling behavior and check the OpenShift driver-container configuration. Reproduce the late-GPU attach and verify the firmware label and workload AVC; done means a late-attached GPU initializes without a custom broad SELinux policy.
Written by the indexing model from the issue text.
Description
On OpenShift with SELinux enforcing, a GPU whose first initialization happens after the driver container has finished installing cannot be initialized for compute. The kernel fails to load gsp_tu10x.bin with error -13 (EACCES). The cause is an SELinux denial: the workload process (container_t) is not allowed to read the firmware file, which is labeled container_var_run_t.
The rest of the driver rootfs under /run/nvidia/driver is labeled container_file_t, which workloads can read. Only the firmware directory has a different label.
The GSP firmware is read at first GPU initialization, by whichever process first accesses the device. This matters because there are two paths:
- Privileged initialization path: a GPU that is already present when the driver container installs is initialized by
nvidia-installerandnvidia-persistenced, both running asspc_t. The firmware label does not matter to them. - Workload initialization path: a GPU whose first initialization happens after the driver installation is first accessed by the workload. The workload is
container_tand is denied.
In our environment, the second GPU becomes available to the node only after the driver container has started. That is how we get a GPU on the workload initialization path.
Visible symptoms: the GPU shows up in nvidia-smi in the driver pod with 0 MiB used and no processes. Node Feature Discovery and GFD still see and label it. The workload does not crash on its own.
To Reproduce
- OpenShift 4.22.9 worker, SELinux enforcing, GPU Operator v26.3.3. One GPU (GPU A) is present when the driver container installs.
- Start a workload that uses GPU A. It works. First GPU initialization of GPU A was done by the privileged initialization path.
- After the driver installation has completed, a second GPU (GPU B) becomes available to the node. How this happens is specific to our environment.
- Start a second workload that requests GPU B. The first device access to GPU B comes from the workload.
- The workload starts, but GPU B is not usable. The kernel log shows the firmware load failing, and the audit log shows the AVC denial for
pythonincontainer_t.
We have not tried a case without a late-appearing GPU (see "Not confirmed yet").
Expected behavior
A GPU whose first initialization happens after driver installation can be initialized and used by a normal workload, with no extra SELinux policy. For example, the firmware files carry a label that container_t can read, or a privileged component performs the first GPU initialization.
Environment
- GPU Operator Version: v26.3.3
- NVIDIA DRA driver for GPUs: v0.4.1
- NVIDIA driver: 580.126.20 (open kernel module), GPU: NVIDIA A30
- OS: RHCOS (OpenShift 4.22.9)
- Kernel Version: 5.14.0-687.35.1.el9_8.x86_64
- Container Runtime Version: CRI-O 1.35.5
- Kubernetes Distro and Version: OpenShift (RHOCP) 4.22.9
Evidence
The SELinux workaround module was removed before the run.
1. The workload is denied on the firmware file
type=SYSCALL msg=audit(...): syscall=257 success=no exit=-5
comm="python" exe="/usr/bin/python3.12"
subj=system_u:system_r:container_t:s0:c7,c28
type=AVC msg=audit(...): avc: denied { read }
comm="python" name="gsp_tu10x.bin" dev="tmpfs" ino=14990
scontext=system_u:system_r:container_t:s0:c7,c28
tcontext=system_u:object_r:container_var_run_t:s0 tclass=file permissive=0
The same denial repeats a few milliseconds later. GPU A, initialized on the privileged path, produced zero AVC denials in the same run. The device node itself is readable: /dev/nvidia1 carries container_file_t:s0:c7,c28.
2. Kernel log, same second as the denial
kernel: nvidia 0000:96:00.0: loading /run/nvidia/driver/lib/firmware/nvidia/580.126.20/gsp_tu10x.bin failed with error -13
kernel: nvidia 0000:96:00.0: Direct firmware load for nvidia/580.126.20/gsp_tu10x.bin failed with error -2
kernel: NVRM: RmFetchGspRmImages: No firmware image found
kernel: NVRM: GPU 0000:96:00.0: RmInitAdapter failed! (0x61:0x56:1914)
kernel: NVRM: GPU 0000:96:00.0: rm_init_adapter failed, device minor number 1
-13 is EACCES. -2 is the fallback path (ENOENT). There is no equivalent failure for GPU A (0000:95:00.0). GPU B is 0000:96:00.0.
3. The privileged path succeeds on the same files
Captured with auditctl -w on the firmware directory (logs allowed and denied accesses):
11:53:26 comm="nvidia-installe" subj=system_u:system_r:spc_t:s0 success=yes (creates gsp_tu10x.bin, gsp_ga10x.bin)
11:53:31 comm="nvidia-persiste" subj=system_u:system_r:spc_t:s0 success=yes (opens gsp_tu10x.bin)
The same pair, in the same order, was seen in two separate runs.
4. The failure is not specific to one device
In an earlier test (2026-09-15), the first device worked (22379 MiB used, active python process, no denial). Later in that test, a workload's first access to the same device after the driver installation failed with the same denial, on the same file and label:
Tue Sep 15 10:26:18 2026 avc: denied { read } for pid=341212 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
Tue Sep 15 10:50:14 2026 avc: denied { read } for pid=380168 comm="python" name="gsp_tu10x.bin" ... tcontext=...container_var_run_t:s0
So the variable is which path performs the first GPU initialization, not which physical device or slot it is.
5. Label mismatch inside the same mount
$ ls -Z /run/nvidia/driver/lib/firmware/nvidia/580.126.20/
system_u:object_r:container_var_run_t:s0 gsp_ga10x.bin
system_u:object_r:container_var_run_t:s0 gsp_tu10x.bin
$ ls -Z /run/nvidia/driver/lib64/ | head
system_u:object_r:container_file_t:s0 LLVMgold.so
system_u:object_r:container_file_t:s0 Mcrt1.o
...
container_t can read container_file_t under the base policy. It cannot read container_var_run_t. Same result in all three runs. The mount is on tmpfs and does not exist before the driver container installs.
6. Driver modules load once per node
The nvidia, nvidia-modeset and nvidia-nvlink modules load once, during the driver installation. No module load happens when GPU B is first initialized six minutes later. So the GSP firmware read is per device, at first GPU initialization.
Other platforms we tested
Same test on each: one GPU present at driver installation, workload on it, then a second GPU whose first initialization happens after driver installation. FACT = seen in the logs we collected. HYP = our assumption. UNKNOWN = the capture failed or was not taken.
| Area | Ubuntu + containerd | RHEL + containerd | RKE2 (SLES 16) | OpenShift (CRI-O) |
|---|---|---|---|---|
| GPU initialized after driver installation is usable | FACT: yes (79 GB used) | FACT: yes | FACT: yes | FACT: no, RmInitAdapter failed |
| GPU model / driver | A100 80GB / 580.126.20 | A100 80GB / 580.126.20 | A100 80GB / 595.71.05 | A30 / 580.126.20 |
| SELinux | UNKNOWN: not captured | FACT: enforcing | FACT: enabled (driver log). Mode not captured | FACT: enforcing |
| Runtime SELinux support | UNKNOWN | FACT: containerd enableSelinux: true |
UNKNOWN | FACT: enabled |
Denial on gsp_*.bin |
UNKNOWN: no audit data | UNKNOWN: audit watch failed, pipeline audit files empty. A short host AVC sample has no firmware denial | FACT: ausearch returned no matches. UNKNOWN whether it covered the first initialization |
FACT: repeated, reproducible |
Kernel -13 / RmInitAdapter |
UNKNOWN: no kernel log | FACT: none in the host journal we have. It does not cover the test window | FACT: none in the kernel log | FACT: present, same second as the denial |
| Label of the firmware directory | UNKNOWN | UNKNOWN: ls -Z returned "Permission denied" |
UNKNOWN: script used a hardcoded 580.126.20 path, node has 595.71.05 | FACT: container_var_run_t |
Privileged nvidia-smi before the workload |
Not done | Not done | Skipped (NODE not set) |
Not done |
| First firmware read on the second GPU | UNKNOWN | UNKNOWN | UNKNOWN | FACT: the workload (python, container_t) |
What we take from this, and what we do not:
- The GPU initialized after driver installation works on the three other platforms. We do not know why. It could be a different firmware label, or a different process doing the first firmware read.
- The driver container script calls
chcon -R -t container_file_t /run/nvidia/driver/devwhen SELinux is enabled (seen in the RKE2 log, and the same step is printed in the RHEL log). It does not relabel the firmware directory. We have not checked whether the OpenShift driver container uses the same script. - The driver message
knvlinkCoreShutdownDeviceLinks_IMPL ... for GPU1also appears on RKE2 and RHEL runs where everything works. We think it is unrelated.
Confirmed vs. not confirmed
Confirmed by logs on this OpenShift setup:
- The firmware file is labeled
container_var_run_t, the workload is denied, and the kernel firmware load fails with-13at the same second. spc_tcomponents (installer, persistenced) read the same file without problems.- The behavior reproduced in three runs.
- With a local SELinux module that allows
container_tto readcontainer_var_run_t, the same steps show zero denials.
Not confirmed yet:
- Whether the same denial happens for a GPU that is present at driver installation but whose first initialization comes from a workload (for example, with
nvidia-persistencedstopped). This would give a reproduction without a late-appearing GPU. Not tested. - The firmware label and the first firmware read on other platforms (see the table).
- Whether
nvidia-persistencedinitializes devices that become available after its start. - Which component creates
/run/nvidia/driver/lib/firmwareand sets its label inside the driver container.
Workaround (works, not a fix)
A custom policy module allowing container_t to read container_var_run_t files. In a run where it was left installed by mistake, the same steps produced no denials. Illustrative form (our local module may differ in detail):
allow container_t container_var_run_t:file { read open getattr };
This widens access for all container_t workloads, which is why we would like a proper fix.
Possible fixes
- Label: make the firmware directory use
container_file_t, like the rest of/run/nvidia/driver. The driver script already relabels/run/nvidia/driver/devwithchcon, so extending that step to/run/nvidia/driver/lib/firmwaremay be a small change. Not tested. - Initialization: make sure a privileged component performs the first GPU initialization for GPUs whose first initialization happens after the driver installation, so the first firmware read never comes from an unprivileged workload.
Information to attach
- kubernetes pods status:
kubectl get pods -n gpu-operator(attached) - kubernetes daemonset status:
kubectl get ds -n gpu-operator(attached) - pod/ds in error state: not applicable, all pods are Running. The failure is inside the workload's GPU initialization.
-
nvidia-smifrom the driver container:kubectl exec DRIVER_POD_NAME -n gpu-operator -c nvidia-driver-ctr -- nvidia-smi(attached, shows GPU B with 0 MiB and no processes) - containerd logs: not applicable, the runtime is CRI-O. CRI-O logs from the node can be added if needed.
- Audit log from the
auditctl -wwatch (keygsp_access_watch) - Full kernel log for the run (
journalctl -k) -
ls -Zoutput of the firmware directory andlib64
- Dominant language
- Go
- Stars
- 2.9k
- Forks
- 569
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 76
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/gpu-operator
-
[Bug]: Driver upgrade does not evict pods that use nvidia.com/gpu only in a native sidecarPossibly taken A pull request linked to this issue is open or already merged. Openbug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/gpu-operator#3026 ·
Maintainers usually reply within 1 day
-
[Bug]: GPUCluster common name label breaks DRA validator selectorPossibly taken @ajavanma claimed this 17 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
NVIDIA/gpu-operator#2955 · 1 comment ·
Maintainers usually reply within 1 day
-
feature lifecycle/frozen needs-triage
Difficulty 4/5 3-5 days Newbie friendliness 25/100
NVIDIA/gpu-operator#3035 · 1 reaction ·
Maintainers usually reply within 1 day
-
Ensure automated backport commits have verified signaturesPossibly taken @asivanadi0 claimed this 8 days ago. Opengood-first-issue
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/gpu-operator#2997 · 1 comment ·
Maintainers usually reply within 1 day
-
[Bug]: Latest Nvidia GPU Operator v26.7.1 reports large numbers of critical and high CVEs in Trivy scan outputPossibly taken @rahulait claimed this 9 days ago. Openmore-information-needed needs-triage
NVIDIA/gpu-operator#2991 · 3 comments · 1 assignee ·
Maintainers usually reply within 1 day
All issues in NVIDIA/gpu-operator
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
duplication
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
openvibely/openvibely#1443 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 80/100
keyxmakerx/Chronicle#1179 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
michelangelo-ai/michelangelo#2258 ·
Maintainers usually reply within 1 day