[Regression][Jetson AGX Thor] Layer-0 V tensor corruption in v0.9.1; v0.4.0 works correctly on the same JetPack 7.1 system
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Domain
- embedded-iot, machine-learning, performance
Research direction
Reproduce the issue on Jetson AGX Thor with the stated JetPack, TensorRT, and Edge-LLM versions, comparing the normal graph with one exposing ONNX linear_2. Inspect the AttentionPlugin input path and the Myelin-generated kernel_mha behavior, using compute-sanitizer as described. Done means the layer-0 V cache matches the reference without exposing the intermediate tensor.
Written by the indexing model from the issue text.
Description
The same Jetson AGX Thor system running JetPack 7.1 produces
correct results with TensorRT Edge-LLM v0.4.0.
TensorRT Edge-LLM v0.9.1 reproduces the layer-0 V-cache corruption
on the same device and software stack.
This suggests an Edge-LLM regression introduced after v0.4.0,
rather than a general JetPack 7.1 compatibility issue.
I am seeing a reproducible numerical correctness issue on Jetson AGX Thor with TensorRT Edge-LLM v0.9.1.
The issue was initially observed as invalid LLM output, but layer-by-layer validation isolates the first divergence to:
decoder layer 0
V projection → AttentionPlugin → V KV cache
The V projection itself is correct. However, in the full graph, the V cache is severely corrupted.
Most importantly, if the V projection tensor is marked as an additional ONNX graph output, the issue completely disappears.
This suggests a possible TensorRT/Myelin tensor lifetime, buffer reuse, or materialization issue around the V tensor.
Environment
Device: Jetson AGX Thor
Jetson Linux: R38.4
JetPack: 7.1
CUDA: 13.0.48
TensorRT: 10.13.3.9
TensorRT Edge-LLM: v0.9.1
Commit: 7f061f21f0a581ba234a1e233c9315b89d8e47d6
Target: jetson-thor
Architecture artifact: aarch64/sm_110
ENABLE_CUTE_DSL=ALL
The repository and submodules are complete.
The problem is reproducible in FP16, so NVFP4 quantization is not the root cause.
The original Hugging Face/PyTorch model runs correctly on the same Thor device.
First numerical divergence
Layer 0 V cache:
cosine = 0.015065519
max_abs = 13.1431694
mean_abs = 0.255735815
shape = [8, 22, 128]
dtype = FP16 vs FP16
NaN/Inf = 0 / 0
Layer 0 K cache from the same execution is essentially correct:
cosine ≈ 1.0
V projection isolation
Testing layer 0 V projection independently:
TensorRT vs PyTorch
cosine = 0.999999881
max_abs = 0.00390625
mean_abs = 0.000058087
So the V projection weights/GEMM themselves appear correct.
Critical observation
Normal full graph
Without exposing the V intermediate tensor:
V cache vs PyTorch V projection
cosine = 0.015077175
Mark V projection as graph output
If ONNX linear_2 (V projection) is marked as an additional graph output:
TensorRT V projection vs PyTorch:
cosine = 0.999998927
And:
V cache vs TensorRT V projection:
cosine = 0.999999881
max_abs = 0
The corruption completely disappears.
Control experiments:
mark Q output → FAIL
mark K output → FAIL
mark V output → PASS
No model weights, inputs, or operator parameters are changed.
CuTe DSL A/B
I also rebuilt Edge-LLM with:
ENABLE_CUTE_DSL=OFF
and confirmed the CuTe DSL FMHA symbol is absent.
The issue still reproduces:
layer 0 V cosine ≈ 0.01507
Therefore this does not appear to be caused by the CuTe DSL FMHA path.
compute-sanitizer
Running:
compute-sanitizer --tool racecheck
on the non-CuTe-DSL build reports:
4 errors
48 warnings
inside the Myelin-generated kernel_mha path.
I cannot confirm that these reports are directly causal, but they appear consistent with the observed behavior.
Expected behavior
The V cache should numerically agree with the PyTorch reference regardless of whether the intermediate V tensor is marked as a graph output.
Actual behavior
V not exposed:
cosine ≈ 0.015
V exposed as graph output:
cosine ≈ 1.0
This makes the issue look like a TensorRT/Myelin optimization, memory lifetime, or buffer reuse problem before AttentionPlugin consumes the V tensor.
- Dominant language
- Python
- Stars
- 563
- Forks
- 135
- Avg merge
- 14h 13m
- Merged PRs (30d)
- 1
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/TensorRT-Edge-LLM
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NVIDIA/TensorRT-Edge-LLM#211 · 2 comments ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/TensorRT-Edge-LLM#183 · 4 comments · 1 reaction ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/TensorRT-Edge-LLM#165 · 2 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NVIDIA/TensorRT-Edge-LLM#144 · 1 comment ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
NVIDIA/TensorRT-Edge-LLM#125 · 1 comment ·
All issues in NVIDIA/TensorRT-Edge-LLM
Similar issues
-
agent-ready documentation needs-triage
Difficulty 1/5 1-3 hours Newbie friendliness 88/100
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 91/100
-
workflow-status page template still says reusable workflows are "triggered only by workflow_call:" Open
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
instance instance add
Difficulty 1/5 Under an hour Newbie friendliness 72/100
searxng/searx-instances#939 · 1 comment ·
-
area-deployment area-integrations triage:bot-seen
Difficulty 2/5 Half a day Newbie friendliness 86/100