Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

SIGFPE (integer divide-by-zero) in cuBLASLt during cuTensorNet path optimization

Đang mở
#232 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
8/100
Loại issue
Lỗi
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Sôi nổi
Công nghệ
hpc, python
Lĩnh vực
quantum-computing

Hướng nghiên cứu

The crash is a SIGFPE inside the closed-source libcublasLt, reached from cutensorEstimateWorkspaceSize during path optimization, so the report's stack traces and the attached cutensornet_sigfpe_capture_517749_rank01.json are the main artifacts to read. Start by checking whether the attached capture reproduces the fault in a single process on H100. Done would mean a minimal reproducer plus a diagnosis that NVIDIA can act on, and this needs maintainer or NVIDIA-side access rather than a first contribution.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Contraction path optimization in cuTensorNet sometimes kills the process with SIGFPE (integer divide-by-zero).
The fault is always inside cublasLtMatmulAlgoGetHeuristic, called from cutensorEstimateWorkspaceSize, called from
cuTensorNet. It happens before any contraction is executed, so it is not a numerical problem with the tensor data.

When running end-to-end adaptive VQE on large circuits (> 150 qubits) on an H100 cluster, I often hit this crash
after a bit more than 100 gates have been added to the circuit. The 180-qubit runs below crashed during adaptive
iterations 104, 125, 125 and 143, and one run already crashed during iteration 71.

The crash seems to be consistent on H100 GPUs, but I have not been able to trigger it on the other GPUs I have tried
(see Reproduction).

Setup

All runs below were on NVIDIA H100 80GB HBM3 GPUs in x86_64 DGX H100 nodes. As an aside, they used 2-8 MPI ranks with
one GPU each, but MPI does not seem to matter: each rank does its own path optimization, and the standalone reproducer
below crashes in a single process without MPI.

Multiple ways to hit the bug

So far I have hit the bug at four different faulting addresses in cuBLASLt, two on each of two cuTensorNet/cuBLASLt
versions:

# cuTensorNet cuBLASLt Faulting instruction Called from Jobs
1 2.13.0 libcublasLt.so.12 libcublasLt.so.12(+0x1a756dd) cuTensorNet worker thread 180-qubit adaptive VQE (iteration 71), one rank of a 180-qubit adaptive VQE run (iteration 125) whose other crashing ranks hit case 2
2 2.13.0 libcublasLt.so.12 libcublasLt.so.12(+0x3879ca7) cutensornetContractionOptimize on the calling thread, and cuTensorNet worker thread Contracting a saved 180-qubit network (7-8 of 8 ranks crashed), 180-qubit adaptive VQE (iterations 125 and 143)
3 2.14.0 libcublasLt.so.13 libcublasLt.so.13(+0x1be9507) cuTensorNet worker thread 180-qubit adaptive VQE (iteration 104)
4 2.14.0 libcublasLt.so.13 libcublasLt.so.13(+0x3e0d107) cuTensorNet worker thread 180-qubit adaptive VQE (iteration 125), the run the reproducer below was captured from

cutensornet_sigfpe_capture_517749_rank01.json

In all four cases the stack is cublasLtMatmulAlgoGetHeuristic <- cutensorEstimateWorkspaceSize+0x527 <-
libcutensornet.so.2.

1. cuTensorNet 2.13.0, libcublasLt.so.12(+0x1a756dd)
[dgx041:2215035] *** Process received signal ***
[dgx041:2215035] Signal: Floating point exception (8)
[dgx041:2215035] Signal code: Integer divide-by-zero (1)
[dgx041:2215035] Failing at address: 0x78f5bd8756dd
[dgx041:2215035] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x78f86d845330]
[dgx041:2215035] [ 1] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1a756dd)[0x78f5bd8756dd]
[dgx041:2215035] [ 2] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1018c2c)[0x78f5bce18c2c]
[dgx041:2215035] [ 3] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x104c96d)[0x78f5bce4c96d]
[dgx041:2215035] [ 4] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x110690d)[0x78f5bcf0690d]
[dgx041:2215035] [ 5] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1106bb0)[0x78f5bcf06bb0]
[dgx041:2215035] [ 6] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(cublasLtMatmulAlgoGetHeuristic+0xa33)[0x78f5bcf3ac83]
[dgx041:2215035] [ 7] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6e9cf5)[0x78f53a6e9cf5]
[dgx041:2215035] [ 8] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ea7a9)[0x78f53a6ea7a9]
[dgx041:2215035] [ 9] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x413f0b)[0x78f53a413f0b]
[dgx041:2215035] [10] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x414053)[0x78f53a414053]
[dgx041:2215035] [11] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x445720)[0x78f53a445720]
[dgx041:2215035] [12] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x44e42c)[0x78f53a44e42c]
[dgx041:2215035] [13] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x4374bc)[0x78f53a4374bc]
[dgx041:2215035] [14] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x78f53a68d5f7]
[dgx041:2215035] [15] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x184cf1)[0x78f533384cf1]
[dgx041:2215035] [16] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1856a7)[0x78f5333856a7]
[dgx041:2215035] [17] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1947fa)[0x78f5333947fa]
[dgx041:2215035] [18] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1640a4)[0x78f5333640a4]
[dgx041:2215035] [19] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x140622)[0x78f533340622]
[dgx041:2215035] [20] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x100d28)[0x78f533300d28]
[dgx041:2215035] [21] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0xfe90f)[0x78f5332fe90f]
[dgx041:2215035] [22] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x10a7da)[0x78f53330a7da]
[dgx041:2215035] [23] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x447ae3)[0x78f533647ae3]
[dgx041:2215035] [24] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x9caa4)[0x78f86d89caa4]
[dgx041:2215035] [25] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x129c6c)[0x78f86d929c6c]
[dgx041:2215035] *** End of error message ***
2. cuTensorNet 2.13.0, libcublasLt.so.12(+0x3879ca7)

Here the crash is reached through the public cutensornetContractionOptimize call:

[dgx152:4054661] *** Process received signal ***
[dgx152:4054661] Signal: Floating point exception (8)
[dgx152:4054661] Signal code: Integer divide-by-zero (1)
[dgx152:4054661] Failing at address: 0x782ea3c79ca7
[dgx152:4054661] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x783152445330]
[dgx152:4054661] [ 1] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3879ca7)[0x782ea3c79ca7]
[dgx152:4054661] [ 2] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3851cda)[0x782ea3c51cda]
[dgx152:4054661] [ 3] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3d17081)[0x782ea4117081]
[dgx152:4054661] [ 4] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3d617ec)[0x782ea41617ec]
[dgx152:4054661] [ 5] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0xfd482b)[0x782ea13d482b]
[dgx152:4054661] [ 6] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0xfd52ef)[0x782ea13d52ef]
[dgx152:4054661] [ 7] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1018e8b)[0x782ea1418e8b]
[dgx152:4054661] [ 8] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x104c96d)[0x782ea144c96d]
[dgx152:4054661] [ 9] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x110690d)[0x782ea150690d]
[dgx152:4054661] [10] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1106bb0)[0x782ea1506bb0]
[dgx152:4054661] [11] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(cublasLtMatmulAlgoGetHeuristic+0xa33)[0x782ea153ac83]
[dgx152:4054661] [12] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6e9cf5)[0x782e1ece9cf5]
[dgx152:4054661] [13] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ea7a9)[0x782e1ecea7a9]
[dgx152:4054661] [14] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x413f0b)[0x782e1ea13f0b]
[dgx152:4054661] [15] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x414053)[0x782e1ea14053]
[dgx152:4054661] [16] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x445720)[0x782e1ea45720]
[dgx152:4054661] [17] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x44e42c)[0x782e1ea4e42c]
[dgx152:4054661] [18] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x4374bc)[0x782e1ea374bc]
[dgx152:4054661] [19] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x782e1ec8d5f7]
[dgx152:4054661] [20] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x184cf1)[0x782e17984cf1]
[dgx152:4054661] [21] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1856a7)[0x782e179856a7]
[dgx152:4054661] [22] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1947fa)[0x782e179947fa]
[dgx152:4054661] [23] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1640a4)[0x782e179640a4]
[dgx152:4054661] [24] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x140622)[0x782e17940622]
[dgx152:4054661] [25] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x104231)[0x782e17904231]
[dgx152:4054661] [26] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x107d8d)[0x782e17907d8d]
[dgx152:4054661] [27] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(cutensornetContractionOptimize+0x2f5)[0x782e17908105]
[dgx152:4054661] [28] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/bindings/cycutensornet.cpython-312-x86_64-linux-gnu.so(+0x5eaa)[0x782e52861eaa]
[dgx152:4054661] [29] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/bindings/cutensornet.cpython-312-x86_64-linux-gnu.so(+0x4cc2a)[0x782e0a65fc2a]
[dgx152:4054661] *** End of error message ***
3. cuTensorNet 2.14.0, libcublasLt.so.13(+0x1be9507)
[dgx055:1935142] *** Process received signal ***
[dgx055:1935142] Signal: Floating point exception (8)
[dgx055:1935142] Signal code: Integer divide-by-zero (1)
[dgx055:1935142] Failing at address: 0x71fa399e9507
[dgx055:1935142] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x71fcdf845330]
[dgx055:1935142] [ 1] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1be9507)[0x71fa399e9507]
[dgx055:1935142] [ 2] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19b7242)[0x71fa397b7242]
[dgx055:1935142] [ 3] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1afe949)[0x71fa398fe949]
[dgx055:1935142] [ 4] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff2e8)[0x71fa398ff2e8]
[dgx055:1935142] [ 5] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff8f4)[0x71fa398ff8f4]
[dgx055:1935142] [ 6] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(cublasLtMatmulAlgoGetHeuristic+0x5d9)[0x71fa3995d149]
[dgx055:1935142] [ 7] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ccdc5)[0x71f9e56ccdc5]
[dgx055:1935142] [ 8] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6cd879)[0x71f9e56cd879]
[dgx055:1935142] [ 9] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3ea8fb)[0x71f9e53ea8fb]
[dgx055:1935142] [10] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3eaa43)[0x71f9e53eaa43]
[dgx055:1935142] [11] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x42a185)[0x71f9e542a185]
[dgx055:1935142] [12] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x40cd9d)[0x71f9e540cd9d]
[dgx055:1935142] [13] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x71f9e5669c87]
[dgx055:1935142] [14] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201439)[0x71f9d4a01439]
[dgx055:1935142] [15] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201e5d)[0x71f9d4a01e5d]
[dgx055:1935142] [16] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x210bac)[0x71f9d4a10bac]
[dgx055:1935142] [17] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1dc1a4)[0x71f9d49dc1a4]
[dgx055:1935142] [18] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1b7af2)[0x71f9d49b7af2]
[dgx055:1935142] [19] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1778f8)[0x71f9d49778f8]
[dgx055:1935142] [20] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x17567f)[0x71f9d497567f]
[dgx055:1935142] [21] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x180b2a)[0x71f9d4980b2a]
[dgx055:1935142] [22] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x50ca63)[0x71f9d4d0ca63]
[dgx055:1935142] [23] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x9cb84)[0x71fcdf89cb84]
[dgx055:1935142] [24] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x129ecc)[0x71fcdf929ecc]
[dgx055:1935142] *** End of error message ***
4. cuTensorNet 2.14.0, libcublasLt.so.13(+0x3e0d107)
[dgx154:480398] *** Process received signal ***
[dgx154:480398] Signal: Floating point exception (8)
[dgx154:480398] Signal code: Integer divide-by-zero (1)
[dgx154:480398] Failing at address: 0x718beda0d107
[dgx154:480398] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x718e89445330]
[dgx154:480398] [ 1] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x3e0d107)[0x718beda0d107]
[dgx154:480398] [ 2] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x3de1c0a)[0x718bed9e1c0a]
[dgx154:480398] [ 3] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x431ac21)[0x718bedf1ac21]
[dgx154:480398] [ 4] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x436a878)[0x718bedf6a878]
[dgx154:480398] [ 5] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19a2c3b)[0x718beb5a2c3b]
[dgx154:480398] [ 6] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19a375f)[0x718beb5a375f]
[dgx154:480398] [ 7] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19b756c)[0x718beb5b756c]
[dgx154:480398] [ 8] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1afe949)[0x718beb6fe949]
[dgx154:480398] [ 9] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff2e8)[0x718beb6ff2e8]
[dgx154:480398] [10] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff8f4)[0x718beb6ff8f4]
[dgx154:480398] [11] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(cublasLtMatmulAlgoGetHeuristic+0x5d9)[0x718beb75d149]
[dgx154:480398] [12] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ccdc5)[0x718b974ccdc5]
[dgx154:480398] [13] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6cd879)[0x718b974cd879]
[dgx154:480398] [14] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3ea8fb)[0x718b971ea8fb]
[dgx154:480398] [15] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3eaa43)[0x718b971eaa43]
[dgx154:480398] [16] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x42a185)[0x718b9722a185]
[dgx154:480398] [17] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x40cd9d)[0x718b9720cd9d]
[dgx154:480398] [18] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x718b97469c87]
[dgx154:480398] [19] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201439)[0x718b86801439]
[dgx154:480398] [20] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201e5d)[0x718b86801e5d]
[dgx154:480398] [21] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x210bac)[0x718b86810bac]
[dgx154:480398] [22] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1dc1a4)[0x718b867dc1a4]
[dgx154:480398] [23] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1b7af2)[0x718b867b7af2]
[dgx154:480398] [24] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1778f8)[0x718b867778f8]
[dgx154:480398] [25] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x17567f)[0x718b8677567f]
[dgx154:480398] [26] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x180b2a)[0x718b86780b2a]
[dgx154:480398] [27] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x50ca63)[0x718b86b0ca63]
[dgx154:480398] [28] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x9cb84)[0x718e8949cb84]
[dgx154:480398] [29] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x129ecc)[0x718e89529ecc]
[dgx154:480398] *** End of error message ***

Reproduction

The crash does not happen on every run of the cuTensorNet path finder, but it happens often enough that I was able to
make a reproducer script that retries the path search up to 300 times, and it usually crashes within a few searches.
It only crashes on the H100 GPUs I have tried, not on a DGX Spark (GB10) or a laptop with an NVIDIA T550 Laptop GPU.

Attached are:

  • reproduce_cutensornet_sigfpe.py: standalone reproducer.
  • cutensornet_sigfpe_capture_517749_rank01.json: the network that rank 1 crashed on in case 4, captured just before
    the contract_path call. It contains only the mode labels and shapes of the 362 input tensors, the output modes,
    the dtype (float32), the network options (compute type and memory limit) and the optimizer options. There is no
    tensor data.

The script builds the network from zero-filled CuPy tensors and repeatedly calls
Network.contract_path(optimize=..., prepare_contraction=False) on a fresh Network. It only needs numpy, cupy
and cuquantum installed. With the JSON file in the same directory as the script, run:

python reproduce_cutensornet_sigfpe.py

The script runs in a single process without MPI. I ran it with cuquantum 26.9.0 (cuTensorNet 2.14.0), cupy 14.2.0,
CUDA runtime 13.2 and cuBLASLt 13.8 on an NVIDIA H100 80GB HBM3, and it crashed during the third search. On the DGX
Spark (aarch64, also cuTensorNet 2.14.0 and cuBLASLt 13.8) all 300 searches finished without a crash, and it did not
crash on a T550 laptop either.

cublaslt_659495.log

Note that with threads=14 the path search is not deterministic even with seed=42, so the found paths differ between
searches, and the number of searches before the crash can vary.

Logs from the H100 run

I have attached the full cuBLASLt API log of the run as cublaslt_659495.log (CUBLASLT_LOG_LEVEL=5). The stdout and
stderr of the reproduction script are below, with the list of extension modules shortened. The process exited with
code 136.

stdout:
Starting cutensornet_sigfpe on ...
cuquantum 26.9.0, cutensornet 21400, cupy 14.2.0, CUDA runtime 13020, GPU NVIDIA H100 80GB HBM3
[1/300] 13.06 s: num_slices=512, opt_cost=5.1015e+09
[2/300] 20.41 s: num_slices=512, opt_cost=2.8673e+09
Reproducer exited with status 136; cuBLASLt logs are in ...
Last cublasLtMatmulAlgoGetHeuristic call (the failing one if the process crashed):
[2026-10-09 13:18:55][cublasLt][660471][Api][cublasLtMatmulAlgoGetHeuristic] Adesc=[type=R_32F rows=268435456 cols=16777216 ld=268435456] Bdesc=[type=R_32F rows=268435456 cols=16777216 ld=268435456] Cdesc=[type=R_32F rows=268435456 cols=268435456 ld=268435456] Ddesc=[type=R_32F rows=268435456 cols=268435456 ld=268435456] preference=[maxWavesCount=0.0 maxWorkspaceSizeinBytes=18446744073709551615] computeDesc=[computeType=COMPUTE_32F scaleType=R_32F transb=OP_T]
Finished cutensornet_sigfpe at Fri Oct  9 01:18:58 PM CEST 2026

stderr:
Fatal Python error: Floating-point exception

Thread 0x000071fc8c62a080 (most recent call first):
  File "/opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/tensornet/tensor_network.py", line 859 in contract_path
  File "/opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvmath/internal/utils.py", line 559 in inner
  File "/opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvmath/internal/utils.py", line 601 in inner
  File "/workspace/app/reproduce_cutensornet_sigfpe.py", line 78 in _search
  File "/workspace/app/reproduce_cutensornet_sigfpe.py", line 109 in main
  File "/workspace/app/reproduce_cutensornet_sigfpe.py", line 120 in <module>

Extension modules: numpy._core._multiarray_umath, numpy.linalg._umath_linalg, cupy_backends.cuda._softlink, ..., _cyutility, scipy._cyutility, scipy._lib._ccallback_c (total: 212)
srun: error: dgx048: task 0: Exited with exit code 136

The cuBLASLt log contains 1,277 cublasLtMatmulAlgoGetHeuristic calls but only 1,276 heuristicResults lines. The
last call, shown in the stdout above, is the last line of the log and never gets its heuristicResults. Earlier calls
in the same log that returned normally also had dimensions of 268435456, so it does not look like a simple limit on a
single dimension.

Possible explanation

I think the crash happens because one heuristic query during path optimization is for an extremely large matmul, and
a size computed from it overflows to zero inside cuBLASLt. With transb=OP_T, the last call is a GEMM with
m = n = 268435456 = 2^28 and k = 16777216 = 2^24. That gives m * n * k = 2^80, which wraps around to exactly 0 in
64-bit integer arithmetic (as does 2 * m * n * k = 2^81). An integer division by that product, or by a value derived
from it, would give exactly this integer divide-by-zero. By comparison, m * n * k was at most 2^52 in all 1,276
earlier calls that returned normally, and their output matrices had at most 2^43 elements, against 2^56 for the
failing call.

Presumably the path optimizer is evaluating a candidate contraction that would never be executed. A float32 output of
2^56 elements is far beyond any GPU's memory. cuTENSOR still asks cuBLASLt for a workspace estimate for it. I would
expect cuBLASLt (or cuTENSOR/cuTensorNet before calling it) to reject such sizes with an error instead of crashing
the process. I have not verified this explanation inside cuBLASLt, and I do not know why it only shows up on H100.
It could be that the H100 heuristics take a code path that the other GPUs do not.

Ngôn ngữ chính
Jupyter Notebook
Star
500
Fork
102
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của NVIDIA/cuQuantum

Tất cả issue của NVIDIA/cuQuantum

Issue tương tự

Thêm issue về Quantum Computing

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.