[BUG]: tileiras SIGSEGV compiling dynamic MMA loops around a 128x64 panel
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 45/100
Rechercherichtung
Beginne mit dem eigenständigen Python-Reproducer im Issue und dem ausgegebenen TileIR und führe dann den in den Logs bei O1 und O2 gezeigten Befehl tileiras aus. Die Fehlerbehebung ist abgeschlossen, wenn der reduzierte dynamic-MMA-Kernel keinen SIGSEGV mehr verursacht und nicht unterstützte Kombinationen stattdessen eine Compilerdiagnose erzeugen; überprüfe beide Optimierungsstufen mit einem frischen Cache.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
cuTile Python version
1.5.0. The same reproducer also fails with 1.4.0.
CUDA Toolkit version
13.3 (tileiras V13.3.36)
Which installation method does this occur on?
Pip
Describe the bug
tileiras terminates with SIGSEGV while compiling the standalone kernel below
at optimization level 1 or 2. The reduced trigger combines two runtime TF32 MMA
loops, a 128x32 panel transformation, an intervening store, and a runtime
conditional. Compilation fails before the deliberately small tensors execute.
I expected the program to compile. If its live tile/resource combination is not
supported, I expected a compiler diagnostic rather than a native process crash.
This was reduced from a fused 64-column Cholesky superpanel. The original used
two float32[1,32768,32768] operands and a large grid to reach the specialization;
the compiler-only example needs two float32[1,4,4] operands and one program.
Minimum reproducible example
import os
import torch
import cuda.tile as ct
ConstInt = ct.Constant[int]
OPT_LEVEL = int(os.environ.get("CUTILE_OPT_LEVEL", "2"))
@ct.kernel(opt_level=OPT_LEVEL)
def _left_superpanel(a, out, super_step, block: ConstInt, rows_per_tile: ConstInt):
batch_id = 0
panel_width = 2 * block
row_block = ct.bid(0)
work = ct.load(
a,
(batch_id, row_block, 0),
shape=(1, rows_per_tile, panel_width),
padding_mode=ct.PaddingMode.ZERO,
).reshape((rows_per_tile, panel_width))
cross_panel = ct.load(
a,
(batch_id, 0, 0),
shape=(1, block, block),
padding_mode=ct.PaddingMode.ZERO,
).reshape((block, block))
cols = ct.arange(block, dtype=ct.int32)[None, :]
right_block = ct.load(
out,
(batch_id, 0, 0),
shape=(1, panel_width, panel_width),
padding_mode=ct.PaddingMode.ZERO,
).reshape((panel_width, panel_width))
for prior_super in range(super_step):
left_block = ct.load(
out,
(batch_id, 0, prior_super),
shape=(1, rows_per_tile, panel_width),
padding_mode=ct.PaddingMode.ZERO,
).reshape((rows_per_tile, panel_width))
left = left_block.astype(ct.tfloat32)
right = right_block.astype(ct.tfloat32)
work = ct.mma(left, right, work)
first_panel = ct.extract(work, (0, 0), shape=(rows_per_tile, block))
second_panel = ct.extract(work, (0, 1), shape=(rows_per_tile, block))
second_panel = ct.mma(
first_panel.astype(ct.tfloat32),
cross_panel.astype(ct.tfloat32),
second_panel,
)
for p in range(block):
solved = ct.extract(second_panel, (0, p), shape=(rows_per_tile, 1))
second_panel = ct.where(cols == p, solved, second_panel)
ct.store(
out,
(batch_id, row_block, 0),
second_panel.reshape((1, rows_per_tile, block)),
)
if super_step:
next_diagonal = cross_panel
for prior_super in range(super_step):
factor = ct.load(
out,
(batch_id, 0, 0),
shape=(1, block, panel_width),
padding_mode=ct.PaddingMode.ZERO,
).reshape((block, panel_width))
next_diagonal = ct.mma(
factor.astype(ct.tfloat32),
factor.transpose(0, 1).astype(ct.tfloat32),
next_diagonal,
)
ct.store(
out,
(batch_id, 0, 0),
next_diagonal.reshape((1, block, block)),
)
def main():
a = torch.empty((1, 4, 4), device="cuda", dtype=torch.float32)
out = torch.empty_like(a)
ct.launch(
torch.cuda.current_stream(),
(1,),
_left_superpanel,
(a, out, 0, 32, 128),
)
torch.cuda.synchronize()
if __name__ == "__main__":
main()
Run at either original optimization level in a fresh compiler cache. Crash dumps
are disabled only to avoid the separate masking problem in #92.
run=$(mktemp -d)
CUDA_TILE_CACHE_DIR=off \
CUDA_TILE_TEMP_DIR="$run" \
CUDA_TILE_ENABLE_CRASH_DUMP=0 \
CUTILE_OPT_LEVEL=2 \
python repro.py
Set CUTILE_OPT_LEVEL=1 for the other original configuration.
Relevant log output
Both optimization levels report:
subprocess.CalledProcessError: Command '['/usr/local/cuda/bin/tileiras',
'/tmp/.../kernel....bytecode', '-o',
'/tmp/.../kernel....cubin', '--gpu-name', 'sm_120',
'-O2', '--lineinfo']' died with <Signals.SIGSEGV: 11>.
cuda.tile._exception.TileCompilerExecutionError: Return code -11
Unknown location
For the O1 run, the compiler command contains -O1 and has the same result.
Directly invoking tileiras on the emitted input exits 139 at both O1 and O2 and
emits no cubin. Four independent cold processes, two at each level,
reproduced. The same script still fails at O2 when importing cuTile Python 1.5.0.
Environment
OS: Ubuntu 22.04.5 LTS, Linux 6.8.0-90-generic x86_64
GPU: NVIDIA RTX PRO 6000 Blackwell Server Edition, compute capability 12.0
Driver: 580.126.09
CUDA toolkit: 13.3; nvcc 13.3.33; tileiras V13.3.36
Python: 3.13.14
PyTorch: 2.12.0+cu130 (bundled CUDA runtime 13.0)
cuTile Python: 1.5.0; also reproduced on 1.4.0
CPU: AMD EPYC 9355, 16 vCPUs
Other details
The two original source variants were:
- O2;
- its direct O1 variant, which changed only the kernel decorator.
The two sources differ only by changing the decorator from O2 to O1. Their retained
12,226-byte TileIR files are byte-for-byte identical (SHA-256
162fae627522e4be144e959fd1594d0979f133a4b5677f89523f8511ff08928a),
so they are two optimization settings for one compiler input.
The reduced trigger no longer contains Cholesky initialization, factorization,
triangular solves, block-count logic, input-dependent values, or a multi-program
grid. These changes make the reduced form compile:
- remove either runtime MMA loop;
- remove the 32-iteration panel loop;
- remove the store between the two dynamic regions;
- remove the runtime conditional;
- reduce
rows_per_tilefrom 128 to 64; - reduce
blockfrom 32 to 16.
The minimized bytecode compiles when replayed at O0 without --lineinfo, although the unreduced source also crashed at
O0. This suggests the larger original contains a broader trigger, while the small
example isolates a deterministic O1/O2 boundary.
Contributing Guidelines
- I agree to follow cuTile Python's contributing guidelines
- I searched the open bugs and found no duplicate for this report
- Vorherrschende Sprache
- Python
- Sterne
- 2.2k
- Forks
- 155
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus NVIDIA/cutile-python
-
[BUG]: check_dtype_support rejects family-conditional (sm_XXXa) gpu_code targetsEvtl. vergeben @sylvesterkaczmarek hat das vor 30 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
NVIDIA/cutile-python#105 · 2 Kommentare ·
-
nvidia-runners
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 25/100
NVIDIA/cutile-python#108 ·
-
[BUG]: FFT sample launches `Batch` blocks that each process the full batchEvtl. vergeben @AntonOresten hat das vor 137 Tagen übernommen. Offenbug status: needs-triage
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 68/100
NVIDIA/cutile-python#102 ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 68/100
NVIDIA/cutile-python#101 ·
-
[BUG]: tileiras SIGSEGV when occupancy=2 is requested for a 32-wide fused tile kernelEvtl. vergeben @yanght27 hat das vor 50 Tagen übernommen. Offenbug
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 52/100
NVIDIA/cutile-python#96 · 1 Kommentar ·
Alle Issues in NVIDIA/cutile-python
Ähnliche Issues
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
topoteretes/cognee#5647 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 62/100
Sendspin/sendspin-python-cli#291 ·
Maintainer antworten meist innerhalb von 6 Tagen
-
Assertion-shape guard fails on development: vacuous recorded-iteration assertion in the assetLinks batch_get helper testEvtl. vergeben Ein verknüpfter Pull Request ist offen oder bereits gemergt. Offenbug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
awslabs/visual-asset-management-system#414 ·
Maintainer antworten meist innerhalb von 1 Tag
-
bug v1 v2
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
modelcontextprotocol/python-sdk#3670 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
aicell-lab/bioengine#232 ·
Maintainer antworten meist innerhalb von 1 Tag