[BUG]: DLPack tensors may be released before async launch work is finished
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Área
- backend, performance
Línea de trabajo
Comienza en cext/tile_kernel.cpp siguiendo arrayrepr_dlpack() a través de arrayrepr_dlpack_common() hasta launch(), y compara cuándo se elimina la cápsula con cuándo se llama a cuLaunchKernelEx(). Valida el comportamiento del lifetime con un Producer cuyo Deleter libera a su último Owner o devuelve la memoria a un pool. Se considera terminado cuando el tensor gestionado se conserva hasta que el trabajo de launch asíncrono ha superado de forma segura su uso, con cobertura para la ruta genérica de DLPack.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Version
current main checkout (0477a34 locally); this code path appears to date back to the initial public import as well
CUDA Toolkit Version
not pinned yet; this came from source review rather than a runtime repro on a specific toolkit version
Which installation method(s) does this occur on?
Source
Describe the bug
I think the generic DLPack launch path is releasing consumed DLManagedTensor objects too early.
In cext/tile_kernel.cpp, arrayrepr_dlpack_common() renames the capsule to "used_dltensor" and then immediately calls tensor->deleter(tensor). That happens while kernel arguments are still being prepared, before prepare_launch() returns, and before cuLaunchKernelEx() is called.
For generic __dlpack__ objects, arrayrepr_dlpack() also calls __dlpack__(stream=-1), so the producer is explicitly being told not to synchronize.
My understanding of the DLPack ownership/lifetime contract is that once the consumer takes ownership of the capsule, it should keep the managed tensor alive until the consumer is actually done with it. Releasing it during argument parsing looks wrong on its own, and in the async CUDA launch path it seems like it could become a stale-pointer / premature-release problem for producers that rely on the deleter to hold the export alive until consumer work has safely passed.
What made me look twice is that there is already a comment in the code saying this is "technically an incorrect implementation" and suggesting an event-based deferred release after launch.
The control flow I am looking at is:
arrayrepr_dlpack()calls__dlpack__(stream=-1)arrayrepr_dlpack_common()reads the pointer, renames the capsule, and immediately calls the deleter- the actual kernel enqueue happens later in
launch()viacuLaunchKernelEx()
Minimum reproducible example
I do not have a clean runtime repro yet. I found this during code review because the control flow itself looks off:
1. producer returns a DLPack capsule
2. cuTile reads the exported pointer
3. cuTile immediately calls the DLPack deleter
4. the CUDA kernel launch happens afterwards, asynchronously
The repro shape I would expect to fail is a producer whose deleter drops the last owner or returns memory to a pool before the launched kernel has actually finished using the pointer.
Relevant log output
none yet
Full env printout
not available yet
Other/Misc.
Code pointers on current main:
cext/tile_kernel.cpp:arrayrepr_dlpack_common()cext/tile_kernel.cpp:arrayrepr_dlpack()cext/tile_kernel.cpp:launch()
If I am reading the intent correctly, it seems like the managed tensor probably needs to stay alive until launch work on the relevant stream is safely past the point where the producer can release it, rather than being deleted during argument extraction.
Happy to help with a repro if that would be useful.
- Lenguaje dominante
- Python
- Estrellas
- 2.2k
- Forks
- 155
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA/cutile-python
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
NVIDIA/cutile-python#105 · 2 comentarios ·
-
bug status: needs-triage
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
NVIDIA/cutile-python#102 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 68/100
NVIDIA/cutile-python#101 ·
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
NVIDIA/cutile-python#97 · 1 comentario ·
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
NVIDIA/cutile-python#96 · 1 comentario ·
Todos los issues de NVIDIA/cutile-python
Issues similares
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
canonical/paas-charm#368 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
-
tech debt
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
-
addition to tracking list Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
StevenBlack/hosts#3256 ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
qualcomm/qai-appbuilder#275 ·