[BUG]: Program.compile("ptx") fails on machines without a CUDA driver, although NVRTC does not need one
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 2/5
- Geschätzter Aufwand
- 1-3 Stunden
- Anfängerfreundlichkeit
- 70/100
Rechercherichtung
Beginne in cuda/core/_program.pyx bei _can_load_generated_ptx() (etwa Zeile 815) und dem cache-hit-Zweig von Program.compile (~Zeile 1192); der auslösende Aufruf ist driver_version() in cuda/core/_utils/version.pyx, der libcuda benötigt. Fange den DynamicLibNotFoundError ab und überspringe die Warnung zur Ladbarkeit, anstatt fehlzuschlagen, analog zum cubin-Pfad. Reproduziere mit dem Snippet auf einer Maschine ohne Treiber und überprüfe die vorhandenen cuda.core-Programmtests; fertig bedeutet, dass PTX ohne einen Treiber kompiliert wird und eine Warnung (oder ein stilles Überspringen) den Fehler ersetzt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Is this a duplicate?
- I confirmed there appear to be no duplicate issues for this bug and that I agree to the Code of Conduct
Type of Bug
Runtime Error
Component
cuda.core
Describe the bug
On a Linux machine with no NVIDIA driver installed (a GitHub Actions ubuntu runner), compiling a Program to PTX with the NVRTC backend raises DynamicLibNotFoundError for libcuda.so.1. Compiling the same source to cubin with an explicit arch works on the same machine, so NVRTC itself is available.
The failure comes from the loadability check that runs before PTX compilation. _can_load_generated_ptx() calls driver_version(), which calls cuDriverGetVersion and needs libcuda. The check only decides whether to emit a RuntimeWarning, but when the driver library is missing it raises instead, and the whole compile fails. The same check is on the cache hit path in Program.compile.
This matters for CI jobs that compile kernels to PTX without a GPU, for example to inspect the generated PTX in tests. My workaround was to call cuda.bindings.nvrtc directly with the flags from ProgramOptions.as_bytes("nvrtc", "ptx"), which works without a driver.
How to Reproduce
On a machine without an NVIDIA driver. I used the ubuntu latest runner on GitHub Actions with Python 3.12.
pip install "cuda-core[cu12]==1.2.1"
from cuda.core import Program, ProgramOptions
src = 'extern "C" __global__ void k(float* x) { x[0] = 2.0f * x[0] + 1.0f; }'
# Works without a driver
Program(src, code_type="c++", options=ProgramOptions(arch="sm_75")).compile("cubin")
# Fails without a driver
Program(src, code_type="c++", options=ProgramOptions(arch="compute_75")).compile("ptx")
A public run on a runner with no libcuda shows both cases, the cubin compile passes and the ptx compile fails. https://github.com/VolodymyrLinuxovich/heat-risk-cuda/actions/runs/37217831804
Trimmed traceback from the failing run
cuda/core/_program.pyx:222: in cuda.core._program.Program.compile
cuda/core/_program.pyx:738: in cuda.core._program._program_compile_uncached
cuda/core/_program.pyx:1192: in cuda.core._program.Program_compile
cuda/core/_program.pyx:815: in cuda.core._program._can_load_generated_ptx
cuda/core/_utils/version.pyx:37: in cuda.core._utils.version.driver_version
cuda/bindings/driver.pyx:23154: in cuda.bindings.driver.cuDriverGetVersion
...
cuda.pathfinder._dynamic_libs.load_dl_common.DynamicLibNotFoundError: "cuda" is an NVIDIA driver library and can only be found via system search. Ensure the NVIDIA display driver is installed.
The _can_load_generated_ptx code on main looks the same as in 1.2.1, so I expect main behaves the same way, but I have only run 1.2.1.
Expected behavior
PTX compilation succeeds without a driver, the same as cubin compilation. When the driver version cannot be determined, the loadability check could skip the warning (or warn that loadability is unknown) instead of failing the compile.
I would be glad to open a PR for this if that approach sounds right to you.
Operating System
Ubuntu (GitHub Actions ubuntu latest runner), no NVIDIA driver
nvidia-smi output
Not available, no driver on that machine.
- Vorherrschende Sprache
- Cython
- Sterne
- 3.4k
- Forks
- 334
- Ø Merge
- 1 T. 17 Std.
- Gemergte PRs (30 T.)
- 123
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus NVIDIA/cuda-python
-
triage
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
NVIDIA/cuda-python#2952 ·
Maintainer antworten meist innerhalb von 1 Tag
-
[DOC]: cuda.core 1.1.1 note misstates program cache permissionsEvtl. vergeben @leofang hat das vor 11 Tagen übernommen. Offendocumentation P1
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 88/100
NVIDIA/cuda-python#2717 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
[DOC]: `PinnedMemoryResource.allocate` documents no parametersEvtl. vergeben @Andy-Jost hat das vor 11 Tagen übernommen. Offencuda.core documentation P1
Schwierigkeit 1/5 1-3 Stunden Anfängerfreundlichkeit 90/100
NVIDIA/cuda-python#2712 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
[BUG]: LocatedHeaderDir is mutable, so callers can poison the cached header-directory lookupEvtl. vergeben @rwgk hat das vor 11 Tagen übernommen. Offentriage
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
NVIDIA/cuda-python#2646 · 1 Reaktion ·
Maintainer antworten meist innerhalb von 1 Tag
-
[FEA]: Support inheritance from BufferEvtl. vergeben @leofang hat das vor 11 Tagen übernommen. Offencuda.core triage
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 62/100
NVIDIA/cuda-python#2435 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in NVIDIA/cuda-python
Ähnliche Issues
-
vxc prints a debug line '[flat-codegen] emitted module via the flat path' on every compileEvtl. vergeben @YodHeVauHe hat das heute übernommen. Offendevex good first issue
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
kmmbvnr/rank#196 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 74/100
SciML/ModelingToolkit.jl#5255 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
NVIDIA/cuda-quantum#5539 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
ocaml/ocaml#15132 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag