[Vulkan] LLM decode ~12% slower than XNNPACK on Adreno
I maintainer di solito rispondono entro 1 giorno
@digantdesai ci sta già lavorando.
Dal 21/9/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Vulkan LLM decode is ~11-13% slower than XNNPACK on Adreno 840 (Galaxy S26 Ultra), while prefill is ~2x faster. Qwen3, llama_main, cold-device medians, arms interleaved.
| Vulkan (4w) | XNNPACK (8da4w) | |
|---|---|---|
| 0.6B decode | 98.0 tok/s | 109.0 tok/s |
| 1.7B decode | 41.3 tok/s | 47.6 tok/s |
| 0.6B prefill, 41-token prompt | 1568 tok/s | 818 tok/s |
At 0.6B, decode reads 335 MB of weights per token. The q4gsw GEMV is ~74% of GPU time at ~44.5 GB/s, and the GPU is ~100% busy.
#22941 (fused QK+softmax on the SDPA decode path) is the one change that helped: +2.5-3.9% decode.
Tried, no gain:
- 8da4w (
linear_dq8ca_q4gsw): 22% slower than 4w - one command buffer per token: no change
- fusing SwiGLU (sigmoid + 2 muls): 0.998x
- tin GEMM at M=1 instead of the coop GEMV: 5.3x slower
- int3 weights: would beat XNNPACK, but +40% perplexity even with HQQ+AWQ
Possible directions: cut the non-GEMV work (SDPA ~12%, rms_norm ~5%, no-op view copies ~5%), or a sub-4-bit linear kernel paired with a quantizer that holds accuracy.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin
- Lingua principale
- Python
- Stelle
- 5k
- Fork
- 1.2k
- Merge medio
- 2g 5h
- PR unite (30g)
- 491
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di pytorch/executorch
-
enhancement triaged
Difficoltà 2/5 Mezza giornata Idoneità per principianti 68/100
pytorch/executorch#21640 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement module: examples
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
pytorch/executorch#23164 · 7 commenti · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channelsForse già presa @psiddh l’ha presa 3 giorni fa. Apertamodule: qnn partner: qualcomm
pytorch/executorch#23160 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
[cpu kernels] native_layer_norm: layer_norm_scalar returns NaN on large-mean rows; Half/BF16 at N>=256 slow after #23153Forse già presa @JakeStevens l’ha presa 3 giorni fa. Apertamodule: kernels
pytorch/executorch#23159 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
module: vulkan
Difficoltà 3/5 1-2 giorni Idoneità per principianti 66/100
pytorch/executorch#23158 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di pytorch/executorch
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-2 giorni Idoneità per principianti 70/100
-
FingerprintSplitter raises ZeroDivisionError when int(frac_train * len(dataset)) floors to zeroAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 7 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
lmstudio-ai/mlx-engine#376 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
pyiron/bagofholding#166 ·