Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Bug] v0.10.0: prebuilt sm_87 (Jetson Orin) CuteDSL artifact missing fmha_v2_d* kernel modules, breaks all LLM engine builds

Aperta
#178 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
65/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
cmake, cpp

Direzione di ricerca

Iniziare da cpp/kernels/contextAttentionKernels/cuteDslFMHAV2Runner.h e ispezionare i contenuti di aarch64/sm_87 in cpp/kernels/cuteDSLArtifact, quindi riprodurre il risultato usando il comando CMake documentato su Jetson Orin. Confrontare i moduli fmha_v2 dichiarati con gli artefatti disponibili; il lavoro è completato quando i moduli richiesti sono presenti e la compilazione di edgellmCore per sm_87 termina correttamente.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Describe the bug

v0.10.0's CHANGELOG states: "Replaced the legacy embedded-cubin FMHA-v2 backend with CuTe DSL FMHA-v2 and removed the checked-in FMHA-v2 cubin artifacts." The new cpp/kernels/contextAttentionKernels/cuteDslFMHAV2Runner.h unconditionally declares LazyKernelModule<fmha_v2_d64_Kernel_Module_t> (and d128/d256/d512/paged/sw/bidirectional variants), but the prebuilt cpp/kernels/cuteDSLArtifact/aarch64/sm_87/ artifact checked into the repo at this same commit does not contain any fmha_v2_* headers (confirmed via find cpp/kernels/cuteDSLArtifact/aarch64/sm_87 -iname '*fmha*' → no results; only gemm/gdn/moe/int4_fp16_gemm/ffpa families are present). "fmha is always linked" per cpp/CMakeLists.txt comment, so this is not optional — it blocks compilation of edgellmCore entirely on Orin.

Since context-attention FMHA is required for essentially every LLM engine build, this appears to break llm_build for all models on native Jetson Orin (sm_87) at v0.10.0, not just spec-decode paths.

Steps to reproduce
# On Jetson Orin, JetPack 7.2, CUDA 13.2
git clone https://github.com/NVIDIA/TensorRT-Edge-LLM.git
cd TensorRT-Edge-LLM
git checkout v0.10.0
git submodule update --init --recursive
mkdir build && cd build
cmake .. -DCMAKE_BUILD_TYPE=Release -DTRT_PACKAGE_DIR=/usr \
  -DCMAKE_TOOLCHAIN_FILE=cmake/aarch64_linux_toolchain.cmake \
  -DEMBEDDED_TARGET=jetson-orin -DCUDA_CTK_VERSION=13.2 -DENABLE_CUTE_DSL=ALL
make -j$(nproc)
Actual behavior
cpp/kernels/contextAttentionKernels/cuteDslFMHAV2Runner.h:127:37: error: ‘fmha_v2_d64_Kernel_Module_t’ was not declared in this scope
  127 |     static detail::LazyKernelModule<fmha_v2_d64_Kernel_Module_t> sLLM_d64;
      |                                     ^~~~~~~~~~~~~~~~~~~~~~~~~~~
compilation terminated due to -Wfatal-errors.
make[2]: *** [cpp/CMakeFiles/edgellmCore.dir/build.make:216: cpp/CMakeFiles/edgellmCore.dir/kernels/contextAttentionKernels/cuteDslFMHAV2Runner.cpp.o] Error 1
Expected behavior

cpp/kernels/cuteDSLArtifact/aarch64/sm_87/ should ship the fmha_v2_d* CuteDSL kernel modules matching what the v0.10.0 header requires, same as (presumably) the x86 sm_110/sm_121 artifacts do.

System information (Edge Device)

  • Platform: NVIDIA Jetson Orin NX 16GB (Seeed reComputer J4012)
  • Software release: JetPack 7.2, CUDA 13.2
  • CPU architecture: aarch64
  • GPU compute capability: SM87
  • Build type: Release
  • TensorRT Edge-LLM version: v0.10.0 (71dd1bae032e70771265917ec74d3ff4cad07a10)
  • CMake options: -DEMBEDDED_TARGET=jetson-orin -DCUDA_CTK_VERSION=13.2 -DENABLE_CUTE_DSL=ALL (exact platform-recommended command from the installation docs)
Lingua principale
Python
Stelle
563
Fork
135
Merge medio
14h 13m
PR unite (30g)
1

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/TensorRT-Edge-LLM

Tutte le issue di NVIDIA/TensorRT-Edge-LLM

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.