SIGFPE (integer divide-by-zero) in cuBLASLt during cuTensorNet path optimization
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 8/100
- Tipo di issue
- Bug
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- hpc, python
- Ambito
- quantum-computing
Direzione di ricerca
The crash is a SIGFPE inside the closed-source libcublasLt, reached from cutensorEstimateWorkspaceSize during path optimization, so the report's stack traces and the attached cutensornet_sigfpe_capture_517749_rank01.json are the main artifacts to read. Start by checking whether the attached capture reproduces the fault in a single process on H100. Done would mean a minimal reproducer plus a diagnosis that NVIDIA can act on, and this needs maintainer or NVIDIA-side access rather than a first contribution.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Contraction path optimization in cuTensorNet sometimes kills the process with SIGFPE (integer divide-by-zero).
The fault is always inside cublasLtMatmulAlgoGetHeuristic, called from cutensorEstimateWorkspaceSize, called from
cuTensorNet. It happens before any contraction is executed, so it is not a numerical problem with the tensor data.
When running end-to-end adaptive VQE on large circuits (> 150 qubits) on an H100 cluster, I often hit this crash
after a bit more than 100 gates have been added to the circuit. The 180-qubit runs below crashed during adaptive
iterations 104, 125, 125 and 143, and one run already crashed during iteration 71.
The crash seems to be consistent on H100 GPUs, but I have not been able to trigger it on the other GPUs I have tried
(see Reproduction).
Setup
All runs below were on NVIDIA H100 80GB HBM3 GPUs in x86_64 DGX H100 nodes. As an aside, they used 2-8 MPI ranks with
one GPU each, but MPI does not seem to matter: each rank does its own path optimization, and the standalone reproducer
below crashes in a single process without MPI.
Multiple ways to hit the bug
So far I have hit the bug at four different faulting addresses in cuBLASLt, two on each of two cuTensorNet/cuBLASLt
versions:
| # | cuTensorNet | cuBLASLt | Faulting instruction | Called from | Jobs |
|---|---|---|---|---|---|
| 1 | 2.13.0 | libcublasLt.so.12 |
libcublasLt.so.12(+0x1a756dd) |
cuTensorNet worker thread | 180-qubit adaptive VQE (iteration 71), one rank of a 180-qubit adaptive VQE run (iteration 125) whose other crashing ranks hit case 2 |
| 2 | 2.13.0 | libcublasLt.so.12 |
libcublasLt.so.12(+0x3879ca7) |
cutensornetContractionOptimize on the calling thread, and cuTensorNet worker thread |
Contracting a saved 180-qubit network (7-8 of 8 ranks crashed), 180-qubit adaptive VQE (iterations 125 and 143) |
| 3 | 2.14.0 | libcublasLt.so.13 |
libcublasLt.so.13(+0x1be9507) |
cuTensorNet worker thread | 180-qubit adaptive VQE (iteration 104) |
| 4 | 2.14.0 | libcublasLt.so.13 |
libcublasLt.so.13(+0x3e0d107) |
cuTensorNet worker thread | 180-qubit adaptive VQE (iteration 125), the run the reproducer below was captured from |
cutensornet_sigfpe_capture_517749_rank01.json
In all four cases the stack is cublasLtMatmulAlgoGetHeuristic <- cutensorEstimateWorkspaceSize+0x527 <-
libcutensornet.so.2.
1. cuTensorNet 2.13.0, libcublasLt.so.12(+0x1a756dd)
[dgx041:2215035] *** Process received signal ***
[dgx041:2215035] Signal: Floating point exception (8)
[dgx041:2215035] Signal code: Integer divide-by-zero (1)
[dgx041:2215035] Failing at address: 0x78f5bd8756dd
[dgx041:2215035] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x78f86d845330]
[dgx041:2215035] [ 1] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1a756dd)[0x78f5bd8756dd]
[dgx041:2215035] [ 2] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1018c2c)[0x78f5bce18c2c]
[dgx041:2215035] [ 3] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x104c96d)[0x78f5bce4c96d]
[dgx041:2215035] [ 4] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x110690d)[0x78f5bcf0690d]
[dgx041:2215035] [ 5] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1106bb0)[0x78f5bcf06bb0]
[dgx041:2215035] [ 6] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(cublasLtMatmulAlgoGetHeuristic+0xa33)[0x78f5bcf3ac83]
[dgx041:2215035] [ 7] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6e9cf5)[0x78f53a6e9cf5]
[dgx041:2215035] [ 8] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ea7a9)[0x78f53a6ea7a9]
[dgx041:2215035] [ 9] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x413f0b)[0x78f53a413f0b]
[dgx041:2215035] [10] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x414053)[0x78f53a414053]
[dgx041:2215035] [11] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x445720)[0x78f53a445720]
[dgx041:2215035] [12] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x44e42c)[0x78f53a44e42c]
[dgx041:2215035] [13] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x4374bc)[0x78f53a4374bc]
[dgx041:2215035] [14] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x78f53a68d5f7]
[dgx041:2215035] [15] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x184cf1)[0x78f533384cf1]
[dgx041:2215035] [16] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1856a7)[0x78f5333856a7]
[dgx041:2215035] [17] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1947fa)[0x78f5333947fa]
[dgx041:2215035] [18] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1640a4)[0x78f5333640a4]
[dgx041:2215035] [19] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x140622)[0x78f533340622]
[dgx041:2215035] [20] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x100d28)[0x78f533300d28]
[dgx041:2215035] [21] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0xfe90f)[0x78f5332fe90f]
[dgx041:2215035] [22] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x10a7da)[0x78f53330a7da]
[dgx041:2215035] [23] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x447ae3)[0x78f533647ae3]
[dgx041:2215035] [24] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x9caa4)[0x78f86d89caa4]
[dgx041:2215035] [25] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x129c6c)[0x78f86d929c6c]
[dgx041:2215035] *** End of error message ***
2. cuTensorNet 2.13.0, libcublasLt.so.12(+0x3879ca7)
Here the crash is reached through the public cutensornetContractionOptimize call:
[dgx152:4054661] *** Process received signal ***
[dgx152:4054661] Signal: Floating point exception (8)
[dgx152:4054661] Signal code: Integer divide-by-zero (1)
[dgx152:4054661] Failing at address: 0x782ea3c79ca7
[dgx152:4054661] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x783152445330]
[dgx152:4054661] [ 1] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3879ca7)[0x782ea3c79ca7]
[dgx152:4054661] [ 2] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3851cda)[0x782ea3c51cda]
[dgx152:4054661] [ 3] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3d17081)[0x782ea4117081]
[dgx152:4054661] [ 4] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x3d617ec)[0x782ea41617ec]
[dgx152:4054661] [ 5] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0xfd482b)[0x782ea13d482b]
[dgx152:4054661] [ 6] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0xfd52ef)[0x782ea13d52ef]
[dgx152:4054661] [ 7] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1018e8b)[0x782ea1418e8b]
[dgx152:4054661] [ 8] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x104c96d)[0x782ea144c96d]
[dgx152:4054661] [ 9] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x110690d)[0x782ea150690d]
[dgx152:4054661] [10] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(+0x1106bb0)[0x782ea1506bb0]
[dgx152:4054661] [11] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cublas/lib/libcublasLt.so.12(cublasLtMatmulAlgoGetHeuristic+0xa33)[0x782ea153ac83]
[dgx152:4054661] [12] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6e9cf5)[0x782e1ece9cf5]
[dgx152:4054661] [13] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ea7a9)[0x782e1ecea7a9]
[dgx152:4054661] [14] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x413f0b)[0x782e1ea13f0b]
[dgx152:4054661] [15] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x414053)[0x782e1ea14053]
[dgx152:4054661] [16] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x445720)[0x782e1ea45720]
[dgx152:4054661] [17] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x44e42c)[0x782e1ea4e42c]
[dgx152:4054661] [18] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x4374bc)[0x782e1ea374bc]
[dgx152:4054661] [19] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x782e1ec8d5f7]
[dgx152:4054661] [20] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x184cf1)[0x782e17984cf1]
[dgx152:4054661] [21] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1856a7)[0x782e179856a7]
[dgx152:4054661] [22] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1947fa)[0x782e179947fa]
[dgx152:4054661] [23] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1640a4)[0x782e179640a4]
[dgx152:4054661] [24] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x140622)[0x782e17940622]
[dgx152:4054661] [25] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x104231)[0x782e17904231]
[dgx152:4054661] [26] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x107d8d)[0x782e17907d8d]
[dgx152:4054661] [27] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(cutensornetContractionOptimize+0x2f5)[0x782e17908105]
[dgx152:4054661] [28] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/bindings/cycutensornet.cpython-312-x86_64-linux-gnu.so(+0x5eaa)[0x782e52861eaa]
[dgx152:4054661] [29] /opt/uv-python/cpython-3.12.14-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/bindings/cutensornet.cpython-312-x86_64-linux-gnu.so(+0x4cc2a)[0x782e0a65fc2a]
[dgx152:4054661] *** End of error message ***
3. cuTensorNet 2.14.0, libcublasLt.so.13(+0x1be9507)
[dgx055:1935142] *** Process received signal ***
[dgx055:1935142] Signal: Floating point exception (8)
[dgx055:1935142] Signal code: Integer divide-by-zero (1)
[dgx055:1935142] Failing at address: 0x71fa399e9507
[dgx055:1935142] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x71fcdf845330]
[dgx055:1935142] [ 1] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1be9507)[0x71fa399e9507]
[dgx055:1935142] [ 2] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19b7242)[0x71fa397b7242]
[dgx055:1935142] [ 3] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1afe949)[0x71fa398fe949]
[dgx055:1935142] [ 4] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff2e8)[0x71fa398ff2e8]
[dgx055:1935142] [ 5] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff8f4)[0x71fa398ff8f4]
[dgx055:1935142] [ 6] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(cublasLtMatmulAlgoGetHeuristic+0x5d9)[0x71fa3995d149]
[dgx055:1935142] [ 7] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ccdc5)[0x71f9e56ccdc5]
[dgx055:1935142] [ 8] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6cd879)[0x71f9e56cd879]
[dgx055:1935142] [ 9] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3ea8fb)[0x71f9e53ea8fb]
[dgx055:1935142] [10] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3eaa43)[0x71f9e53eaa43]
[dgx055:1935142] [11] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x42a185)[0x71f9e542a185]
[dgx055:1935142] [12] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x40cd9d)[0x71f9e540cd9d]
[dgx055:1935142] [13] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x71f9e5669c87]
[dgx055:1935142] [14] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201439)[0x71f9d4a01439]
[dgx055:1935142] [15] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201e5d)[0x71f9d4a01e5d]
[dgx055:1935142] [16] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x210bac)[0x71f9d4a10bac]
[dgx055:1935142] [17] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1dc1a4)[0x71f9d49dc1a4]
[dgx055:1935142] [18] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1b7af2)[0x71f9d49b7af2]
[dgx055:1935142] [19] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1778f8)[0x71f9d49778f8]
[dgx055:1935142] [20] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x17567f)[0x71f9d497567f]
[dgx055:1935142] [21] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x180b2a)[0x71f9d4980b2a]
[dgx055:1935142] [22] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x50ca63)[0x71f9d4d0ca63]
[dgx055:1935142] [23] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x9cb84)[0x71fcdf89cb84]
[dgx055:1935142] [24] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x129ecc)[0x71fcdf929ecc]
[dgx055:1935142] *** End of error message ***
4. cuTensorNet 2.14.0, libcublasLt.so.13(+0x3e0d107)
[dgx154:480398] *** Process received signal ***
[dgx154:480398] Signal: Floating point exception (8)
[dgx154:480398] Signal code: Integer divide-by-zero (1)
[dgx154:480398] Failing at address: 0x718beda0d107
[dgx154:480398] [ 0] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x45330)[0x718e89445330]
[dgx154:480398] [ 1] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x3e0d107)[0x718beda0d107]
[dgx154:480398] [ 2] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x3de1c0a)[0x718bed9e1c0a]
[dgx154:480398] [ 3] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x431ac21)[0x718bedf1ac21]
[dgx154:480398] [ 4] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x436a878)[0x718bedf6a878]
[dgx154:480398] [ 5] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19a2c3b)[0x718beb5a2c3b]
[dgx154:480398] [ 6] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19a375f)[0x718beb5a375f]
[dgx154:480398] [ 7] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x19b756c)[0x718beb5b756c]
[dgx154:480398] [ 8] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1afe949)[0x718beb6fe949]
[dgx154:480398] [ 9] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff2e8)[0x718beb6ff2e8]
[dgx154:480398] [10] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(+0x1aff8f4)[0x718beb6ff8f4]
[dgx154:480398] [11] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvidia/cu13/lib/libcublasLt.so.13(cublasLtMatmulAlgoGetHeuristic+0x5d9)[0x718beb75d149]
[dgx154:480398] [12] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6ccdc5)[0x718b974ccdc5]
[dgx154:480398] [13] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x6cd879)[0x718b974cd879]
[dgx154:480398] [14] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3ea8fb)[0x718b971ea8fb]
[dgx154:480398] [15] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x3eaa43)[0x718b971eaa43]
[dgx154:480398] [16] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x42a185)[0x718b9722a185]
[dgx154:480398] [17] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(+0x40cd9d)[0x718b9720cd9d]
[dgx154:480398] [18] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cutensor/lib/libcutensor.so.2(cutensorEstimateWorkspaceSize+0x527)[0x718b97469c87]
[dgx154:480398] [19] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201439)[0x718b86801439]
[dgx154:480398] [20] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x201e5d)[0x718b86801e5d]
[dgx154:480398] [21] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x210bac)[0x718b86810bac]
[dgx154:480398] [22] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1dc1a4)[0x718b867dc1a4]
[dgx154:480398] [23] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1b7af2)[0x718b867b7af2]
[dgx154:480398] [24] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x1778f8)[0x718b867778f8]
[dgx154:480398] [25] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x17567f)[0x718b8677567f]
[dgx154:480398] [26] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x180b2a)[0x718b86780b2a]
[dgx154:480398] [27] /opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/lib/libcutensornet.so.2(+0x50ca63)[0x718b86b0ca63]
[dgx154:480398] [28] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x9cb84)[0x718e8949cb84]
[dgx154:480398] [29] /usr/lib/x86_64-linux-gnu/libc.so.6(+0x129ecc)[0x718e89529ecc]
[dgx154:480398] *** End of error message ***
Reproduction
The crash does not happen on every run of the cuTensorNet path finder, but it happens often enough that I was able to
make a reproducer script that retries the path search up to 300 times, and it usually crashes within a few searches.
It only crashes on the H100 GPUs I have tried, not on a DGX Spark (GB10) or a laptop with an NVIDIA T550 Laptop GPU.
Attached are:
reproduce_cutensornet_sigfpe.py: standalone reproducer.cutensornet_sigfpe_capture_517749_rank01.json: the network that rank 1 crashed on in case 4, captured just before
thecontract_pathcall. It contains only the mode labels and shapes of the 362 input tensors, the output modes,
the dtype (float32), the network options (compute type and memory limit) and the optimizer options. There is no
tensor data.
The script builds the network from zero-filled CuPy tensors and repeatedly calls
Network.contract_path(optimize=..., prepare_contraction=False) on a fresh Network. It only needs numpy, cupy
and cuquantum installed. With the JSON file in the same directory as the script, run:
python reproduce_cutensornet_sigfpe.py
The script runs in a single process without MPI. I ran it with cuquantum 26.9.0 (cuTensorNet 2.14.0), cupy 14.2.0,
CUDA runtime 13.2 and cuBLASLt 13.8 on an NVIDIA H100 80GB HBM3, and it crashed during the third search. On the DGX
Spark (aarch64, also cuTensorNet 2.14.0 and cuBLASLt 13.8) all 300 searches finished without a crash, and it did not
crash on a T550 laptop either.
Note that with threads=14 the path search is not deterministic even with seed=42, so the found paths differ between
searches, and the number of searches before the crash can vary.
Logs from the H100 run
I have attached the full cuBLASLt API log of the run as cublaslt_659495.log (CUBLASLT_LOG_LEVEL=5). The stdout and
stderr of the reproduction script are below, with the list of extension modules shortened. The process exited with
code 136.
stdout:
Starting cutensornet_sigfpe on ...
cuquantum 26.9.0, cutensornet 21400, cupy 14.2.0, CUDA runtime 13020, GPU NVIDIA H100 80GB HBM3
[1/300] 13.06 s: num_slices=512, opt_cost=5.1015e+09
[2/300] 20.41 s: num_slices=512, opt_cost=2.8673e+09
Reproducer exited with status 136; cuBLASLt logs are in ...
Last cublasLtMatmulAlgoGetHeuristic call (the failing one if the process crashed):
[2026-10-09 13:18:55][cublasLt][660471][Api][cublasLtMatmulAlgoGetHeuristic] Adesc=[type=R_32F rows=268435456 cols=16777216 ld=268435456] Bdesc=[type=R_32F rows=268435456 cols=16777216 ld=268435456] Cdesc=[type=R_32F rows=268435456 cols=268435456 ld=268435456] Ddesc=[type=R_32F rows=268435456 cols=268435456 ld=268435456] preference=[maxWavesCount=0.0 maxWorkspaceSizeinBytes=18446744073709551615] computeDesc=[computeType=COMPUTE_32F scaleType=R_32F transb=OP_T]
Finished cutensornet_sigfpe at Fri Oct 9 01:18:58 PM CEST 2026
stderr:
Fatal Python error: Floating-point exception
Thread 0x000071fc8c62a080 (most recent call first):
File "/opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/cuquantum/tensornet/tensor_network.py", line 859 in contract_path
File "/opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvmath/internal/utils.py", line 559 in inner
File "/opt/uv-python/cpython-3.12.15-linux-x86_64-gnu/lib/python3.12/site-packages/nvmath/internal/utils.py", line 601 in inner
File "/workspace/app/reproduce_cutensornet_sigfpe.py", line 78 in _search
File "/workspace/app/reproduce_cutensornet_sigfpe.py", line 109 in main
File "/workspace/app/reproduce_cutensornet_sigfpe.py", line 120 in <module>
Extension modules: numpy._core._multiarray_umath, numpy.linalg._umath_linalg, cupy_backends.cuda._softlink, ..., _cyutility, scipy._cyutility, scipy._lib._ccallback_c (total: 212)
srun: error: dgx048: task 0: Exited with exit code 136
The cuBLASLt log contains 1,277 cublasLtMatmulAlgoGetHeuristic calls but only 1,276 heuristicResults lines. The
last call, shown in the stdout above, is the last line of the log and never gets its heuristicResults. Earlier calls
in the same log that returned normally also had dimensions of 268435456, so it does not look like a simple limit on a
single dimension.
Possible explanation
I think the crash happens because one heuristic query during path optimization is for an extremely large matmul, and
a size computed from it overflows to zero inside cuBLASLt. With transb=OP_T, the last call is a GEMM with
m = n = 268435456 = 2^28 and k = 16777216 = 2^24. That gives m * n * k = 2^80, which wraps around to exactly 0 in
64-bit integer arithmetic (as does 2 * m * n * k = 2^81). An integer division by that product, or by a value derived
from it, would give exactly this integer divide-by-zero. By comparison, m * n * k was at most 2^52 in all 1,276
earlier calls that returned normally, and their output matrices had at most 2^43 elements, against 2^56 for the
failing call.
Presumably the path optimizer is evaluating a candidate contraction that would never be executed. A float32 output of
2^56 elements is far beyond any GPU's memory. cuTENSOR still asks cuBLASLt for a workspace estimate for it. I would
expect cuBLASLt (or cuTENSOR/cuTensorNet before calling it) to reject such sizes with an error instead of crashing
the process. I have not verified this explanation inside cuBLASLt, and I do not know why it only shows up on H100.
It could be that the H100 heuristics take a code path that the other GPUs do not.
- Lingua principale
- Jupyter Notebook
- Stelle
- 500
- Fork
- 102
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/cuQuantum
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 94/100
-
CUDA 12 cuquantum-appliance image ships a UCX `libucm_cuda.so.0` linked against `libcudart.so.13`Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 25/100
-
`cutensornetTensorSVD` returns `CUTENSORNET_STATUS_INTERNAL_ERROR` with missing host workspaceAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 78/100