[BUG]: Program.compile("ptx") fails on machines without a CUDA driver, although NVRTC does not need one
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
研究方向
從 cuda/core/_program.pyx 中的 _can_load_generated_ptx()(約第 815 行)和 Program.compile 的 cache-hit 分支(約第 1192 行)開始;引發異常的呼叫是 cuda/core/_utils/version.pyx 中的 driver_version(),它需要 libcuda。捕獲 DynamicLibNotFoundError 並跳過可載入性警告而不是失敗,模仿 cubin 路徑。在無驅動程式的機器上使用程式碼片段重現,並檢查現有的 cuda.core 程式測試;完成意味著 PTX 在沒有驅動程式的情況下編譯,並且警告(或靜默跳過)取代錯誤。
由索引模型根據 Issue 內容生成。
描述
Is this a duplicate?
- I confirmed there appear to be no duplicate issues for this bug and that I agree to the Code of Conduct
Type of Bug
Runtime Error
Component
cuda.core
Describe the bug
On a Linux machine with no NVIDIA driver installed (a GitHub Actions ubuntu runner), compiling a Program to PTX with the NVRTC backend raises DynamicLibNotFoundError for libcuda.so.1. Compiling the same source to cubin with an explicit arch works on the same machine, so NVRTC itself is available.
The failure comes from the loadability check that runs before PTX compilation. _can_load_generated_ptx() calls driver_version(), which calls cuDriverGetVersion and needs libcuda. The check only decides whether to emit a RuntimeWarning, but when the driver library is missing it raises instead, and the whole compile fails. The same check is on the cache hit path in Program.compile.
This matters for CI jobs that compile kernels to PTX without a GPU, for example to inspect the generated PTX in tests. My workaround was to call cuda.bindings.nvrtc directly with the flags from ProgramOptions.as_bytes("nvrtc", "ptx"), which works without a driver.
How to Reproduce
On a machine without an NVIDIA driver. I used the ubuntu latest runner on GitHub Actions with Python 3.12.
pip install "cuda-core[cu12]==1.2.1"
from cuda.core import Program, ProgramOptions
src = 'extern "C" __global__ void k(float* x) { x[0] = 2.0f * x[0] + 1.0f; }'
# Works without a driver
Program(src, code_type="c++", options=ProgramOptions(arch="sm_75")).compile("cubin")
# Fails without a driver
Program(src, code_type="c++", options=ProgramOptions(arch="compute_75")).compile("ptx")
A public run on a runner with no libcuda shows both cases, the cubin compile passes and the ptx compile fails. https://github.com/VolodymyrLinuxovich/heat-risk-cuda/actions/runs/37217831804
Trimmed traceback from the failing run
cuda/core/_program.pyx:222: in cuda.core._program.Program.compile
cuda/core/_program.pyx:738: in cuda.core._program._program_compile_uncached
cuda/core/_program.pyx:1192: in cuda.core._program.Program_compile
cuda/core/_program.pyx:815: in cuda.core._program._can_load_generated_ptx
cuda/core/_utils/version.pyx:37: in cuda.core._utils.version.driver_version
cuda/bindings/driver.pyx:23154: in cuda.bindings.driver.cuDriverGetVersion
...
cuda.pathfinder._dynamic_libs.load_dl_common.DynamicLibNotFoundError: "cuda" is an NVIDIA driver library and can only be found via system search. Ensure the NVIDIA display driver is installed.
The _can_load_generated_ptx code on main looks the same as in 1.2.1, so I expect main behaves the same way, but I have only run 1.2.1.
Expected behavior
PTX compilation succeeds without a driver, the same as cubin compilation. When the driver version cannot be determined, the loadability check could skip the warning (or warn that loadability is unknown) instead of failing the compile.
I would be glad to open a PR for this if that approach sounds right to you.
Operating System
Ubuntu (GitHub Actions ubuntu latest runner), no NVIDIA driver
nvidia-smi output
Not available, no driver on that machine.
- 主要語言
- Cython
- 星號
- 3.4k
- 分支
- 335
- 平均合併
- 1 天 18 小時
- 30 天內合併 PR
- 146
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
NVIDIA/cuda-python 的其他 Issue
-
難度 2/5 1-3 小時 新手友好度 84/100
NVIDIA/cuda-python#2952 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[DOC]: cuda.core 1.1.1 note misstates program cache permissions可能已有人在做 @leofang 於 15 天前認領。 未關閉documentation P1
難度 1/5 1 小時以內 新手友好度 88/100
NVIDIA/cuda-python#2717 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[DOC]: `PinnedMemoryResource.allocate` documents no parameters可能已有人在做 @Andy-Jost 於 15 天前認領。 未關閉cuda.core documentation P1
難度 1/5 1-3 小時 新手友好度 90/100
NVIDIA/cuda-python#2712 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[BUG]: LocatedHeaderDir is mutable, so callers can poison the cached header-directory lookup可能已有人在做 @rwgk 於 15 天前認領。 未關閉triage
難度 2/5 1-3 小時 新手友好度 82/100
NVIDIA/cuda-python#2646 · 1 個 reaction ·
維護者通常 1 天內回覆
-
[FEA]: Support inheritance from Buffer可能已有人在做 @leofang 於 15 天前認領。 未關閉cuda.core triage
難度 2/5 1-3 小時 新手友好度 62/100
NVIDIA/cuda-python#2435 · 1 則留言 ·
維護者通常 1 天內回覆
查看 NVIDIA/cuda-python 的全部 Issue
相似的 Issue
-
Default-import note suggests `import * as process` for velt:process, which does not name the builtin未關閉
難度 2/5 1-3 小時 新手友好度 72/100
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 76/100
維護者通常 1 天內回覆
-
The XMIR reader wraps an `@as` index above `Int`, so `as="α18446744073709551616"` is read as `α0`未關閉
難度 2/5 1-3 小時 新手友好度 78/100
objectionary/phino#1769 ·
維護者通常 1 天內回覆
-
A `\u` in an object name breaks the Javadoc that `to-java.xsl` writes, so the class does not compile未關閉
難度 2/5 1-3 小時 新手友好度 74/100
objectionary/eo#9347 ·
維護者通常 1 天內回覆
-
erdos-status-sync
難度 2/5 1-3 小時 新手友好度 65/100
google-deepmind/formal-conjectures#6920 ·
維護者通常 1 天內回覆