NCCL EP hard link makes `import transformer_engine` require `libcuda.so.1`
维护者通常 2 天内回复
@phu0ngng 已经在做这个了。
开始于 2026年8月15日。
评估
这个 Issue 还没有评估数据。
描述
Describe the bug
When Transformer Engine is built for Hopper or newer with NCCL EP enabled,
libtransformer_engine.so has a direct ELF DT_NEEDED dependency on
libcuda.so.1. As a result, plain import transformer_engine fails on a host
that has the CUDA toolkit/runtime libraries but intentionally has no NVIDIA
driver or GPU device.
This happens before any Transformer Engine operation, CUDA execution, or NCCL
EP API is requested. It prevents CPU-only workflows that need to import TE and
construct TE modules with parameters on device="cpu", while never running a
TE forward pass. One downstream example is CPU-only checkpoint conversion:
using the normal TE-backed model spec preserves the exact parameter and
checkpoint schema expected by the GPU model, whereas substituting local
PyTorch modules can change fused parameter names and TE _extra_state entries.
The direct driver dependency was added by
#3127, commit
4955320121e500db98f1c08d8d90075c17d9469e. The NCCL EP CMake path
whole-archives libnccl_ep.a into libtransformer_engine.so and explicitly
links CUDA::cuda_driver. The accompanying comment says this ordering is
intended to make --as-needed record libcuda.so.1.
This regresses the loading property established by
#1240, which removed
the core library's direct CUDA driver link and used TE's existing indirect
driver-entry-point infrastructure instead. The hard link is still present on
TE main at 2d80391f77e542c06d0281688ddd17b6e34adec8.
Relevant source:
- NCCL EP link at the affected revision:
https://github.com/NVIDIA/TransformerEngine/blob/4329ff84bfbdaa778a33cba02a15fb0807c64689/transformer_engine/common/CMakeLists.txt#L440-L510 - Import-time core-library load:
https://github.com/NVIDIA/TransformerEngine/blob/4329ff84bfbdaa778a33cba02a15fb0807c64689/transformer_engine/common/__init__.py#L357-L382 - Current-main NCCL EP link:
https://github.com/NVIDIA/TransformerEngine/blob/2d80391f77e542c06d0281688ddd17b6e34adec8/transformer_engine/common/CMakeLists.txt#L496-L514
Steps/Code to reproduce bug
Use a TE build that targets SM90 or newer and has NCCL EP enabled. Run it on a
Linux host/container with the required CUDA toolkit libraries installed, but
with no libcuda.so.1 and no /dev/nvidia* devices.
$ ls /dev/nvidia*
ls: cannot access '/dev/nvidia*': No such file or directory
$ ldconfig -p | grep -E 'libcuda\.so|libnvidia-ml'
# no output
$ python - <<'PY'
import torch
print("CUDA initialized before TE import:", torch.cuda.is_initialized())
import transformer_engine
PY
CUDA initialized before TE import: False
Traceback (most recent call last):
...
File "transformer_engine/common/__init__.py", line 382, in <module>
_TE_LIB_CTYPES = _load_core_library()
File "transformer_engine/common/__init__.py", line 360, in _load_core_library
return ctypes.CDLL(..., mode=ctypes.RTLD_GLOBAL | os.RTLD_LAZY)
OSError: libcuda.so.1: cannot open shared object file: No such file or directory
The binary dependency is visible without importing TE:
$ readelf -d /path/to/libtransformer_engine.so | grep NEEDED | grep libcuda
0x0000000000000001 (NEEDED) Shared library: [libcuda.so.1]
RTLD_LAZY does not help because the ELF loader must resolve direct
DT_NEEDED libraries when libtransformer_engine.so is loaded.
The observed PyTorch extension did not itself have a direct libcuda.so.1
entry; the failing dependency was on the TE core library.
Expected behavior
On a system where TE's non-driver shared-library dependencies are present,
importing transformer_engine and transformer_engine.pytorch should not
require the NVIDIA driver merely because NCCL EP was included at build time.
Constructing TE modules with parameters on device="cpu" should remain
possible without initializing or using CUDA. CUDA execution and NCCL EP may
still require a driver and GPU, and should fail with a clear error only when
those capabilities are actually requested.
Suggested fix
Prefer making the NCCL EP backend an optional, lazily loaded component instead
of whole-archiving it into the always-loaded TE core library:
- Keep the public
nvte_ep_*C API inlibtransformer_engine.soas thin
forwarding entry points. - Put
ep_backend.cpp,libnccl_ep.a, and their NCCL/CUDA-driver link
dependencies in a separate shared object, for example
libtransformer_engine_nccl_ep.so. - Load that backend with
dlopenand resolve a versioned function table on
the firstnvte_ep_initialize()call, not during Python package import. - If the backend or driver is unavailable, raise an actionable NCCL EP error
at that point. Other TE imports and CPU parameter construction should remain
usable. - Preserve the current throwing stubs when TE is built with
NVTE_WITH_NCCL_EP=0.
An alternative is to remove direct CUDA driver references from the NCCL EP
objects and route them through TE's existing cudaGetDriverEntryPoint-based
loader. The key requirement is that the always-loaded
libtransformer_engine.so no longer records libcuda.so.1 solely because the
optional NCCL EP backend was compiled.
Current workaround
Building TE with NVTE_WITH_NCCL_EP=0 selects the existing throwing
nvte_ep_* stubs and avoids the NCCL EP CMake link branch. This is suitable
for a CPU-only conversion image, but it disables NCCL EP for GPU/MoE workloads
and therefore is not a general solution for a shared training image.
Adding a CUDA stub library to LD_LIBRARY_PATH is not a safe workaround. It
masks the import-time dependency and defers failure until an accidental driver
call.
Proposed acceptance tests
-
Build for SM90+ with
NVTE_WITH_NCCL_EP=1and verify that
libtransformer_engine.sohas noDT_NEEDEDentry forlibcuda.so.1.
A separately loaded NCCL EP backend may retain that dependency. -
In a Linux container with CUDA toolkit/runtime libraries but no NVIDIA
driver or/dev/nvidia*, verify:import torch import transformer_engine import transformer_engine.pytorch as te assert not torch.cuda.is_initialized() module = te.Linear(16, 16, device="cpu") assert module.weight.device.type == "cpu" assert not torch.cuda.is_initialized() -
Verify that requesting NCCL EP without a usable driver/backend produces a
clear runtime exception at the EP API boundary rather than breaking package
import. -
Run the existing NCCL EP tests on H100 or newer and confirm there is no
functional or performance regression after the backend is loaded. -
Retain coverage for
NVTE_WITH_NCCL_EP=0and its existing stubs.
Environment overview
- Environment location: Linux container on a genuinely driverless CPU-only
host - Installation: prebuilt container package from the exact TE source revision
below - Transformer Engine:
2.17.1+4329ff84 - Transformer Engine source:
4329ff84bfbdaa778a33cba02a15fb0807c64689 - Python: 3.12.3
- PyTorch:
2.13.0a0+8145d630e8.nv26.06 - CUDA toolkit/runtime: 13.3
- NVIDIA driver library: absent
- GPU devices: none
Device details
No GPU is intentionally present for the failing import. The TE binary was
built for Hopper-or-newer targets, which enables the NCCL EP build path by
default.
Additional context
This report does not request CPU execution of TE kernels. The required contract
is narrower: importing TE and constructing its parameter schema on CPU should
not load the NVIDIA driver. Normal TE forward execution and NCCL EP remain GPU
operations.
- 主要语言
- Python
- 星标
- 3.6k
- 派生
- 851
- 平均合并
- 5 天 1 小时
- 30 天内合并 PR
- 52
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/TransformerEngine 的其他 Issue
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it未关闭
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA/TransformerEngine#3647 ·
维护者通常 2 天内回复
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar可能已有人在做 @sanjana658 于 1 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA/TransformerEngine#3636 · 2 条评论 ·
维护者通常 2 天内回复
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it可能已有人在做 @yuweih205 于 32 天前认领。 未关闭attention
难度 2/5 1-3 小时 新手友好度 85/100
NVIDIA/TransformerEngine#3481 · 4 条评论 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
NVIDIA/TransformerEngine#2189 · 7 条评论 · 5 个 reaction ·
维护者通常 2 天内回复
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration builds可能已有人在做 @ksivaman 今天认领。 未关闭
难度 4/5 3-5 天 新手友好度 50/100
NVIDIA/TransformerEngine#3645 · 1 条评论 · 已指派 1 人 ·
维护者通常 2 天内回复
查看 NVIDIA/TransformerEngine 的全部 Issue
相似的 Issue
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0未关闭documentation priority: low size: S
难度 2/5 1-3 小时 新手友好度 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
维护者通常 1 天内回复
-
bug help wanted
难度 2/5 1-3 小时 新手友好度 75/100
维护者通常 1 天内回复
-
documentation
难度 1/5 1 小时以内 新手友好度 65/100
ansys/pydpf-core#3547 ·
维护者通常 1 天内回复
-
core
难度 2/5 1-3 小时 新手友好度 70/100
vectorize-io/hindsight#5457 ·
维护者通常 1 天内回复
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recording可能已有人在做 @ktz03 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 72/100
volcengine/OpenViking#5806 ·
维护者通常 1 天内回复