[Vulkan] LLM decode ~12% slower than XNNPACK on Adreno
Maintainer thường phản hồi trong vòng 1 ngày
@digantdesai đang làm issue này rồi.
Từ ngày 21/9/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Vulkan LLM decode is ~11-13% slower than XNNPACK on Adreno 840 (Galaxy S26 Ultra), while prefill is ~2x faster. Qwen3, llama_main, cold-device medians, arms interleaved.
| Vulkan (4w) | XNNPACK (8da4w) | |
|---|---|---|
| 0.6B decode | 98.0 tok/s | 109.0 tok/s |
| 1.7B decode | 41.3 tok/s | 47.6 tok/s |
| 0.6B prefill, 41-token prompt | 1568 tok/s | 818 tok/s |
At 0.6B, decode reads 335 MB of weights per token. The q4gsw GEMV is ~74% of GPU time at ~44.5 GB/s, and the GPU is ~100% busy.
#22941 (fused QK+softmax on the SDPA decode path) is the one change that helped: +2.5-3.9% decode.
Tried, no gain:
- 8da4w (
linear_dq8ca_q4gsw): 22% slower than 4w - one command buffer per token: no change
- fusing SwiGLU (sigmoid + 2 muls): 0.998x
- tin GEMM at M=1 instead of the coop GEMV: 5.3x slower
- int3 weights: would beat XNNPACK, but +40% perplexity even with HQQ+AWQ
Possible directions: cut the non-GEMV work (SDPA ~12%, rms_norm ~5%, no-op view copies ~5%), or a sub-4-bit linear kernel paired with a quantizer that holds accuracy.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin
- Ngôn ngữ chính
- Python
- Star
- 5k
- Fork
- 1.2k
- Merge trung bình
- 2 ngày 5 giờ
- Pull request đã merge (30 ngày)
- 491
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của pytorch/executorch
-
enhancement triaged
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 68/100
pytorch/executorch#21640 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[RFC] ExecuTorch Persisting Device Specialized Delegate ArtifactsCó thể đã có người làm @JacobSzwejbka đã nhận hôm nay. Đang mở
pytorch/executorch#23192 · 1 bình luận · 1 reaction · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[QNN] Enable ConvTranspose + BatchNorm fusion after #23170Có thể đã có người làm @psiddh đã nhận hôm nay. Đang mởmodule: qnn partner: qualcomm
pytorch/executorch#23185 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement module: examples
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 20/100
pytorch/executorch#23164 · 7 bình luận · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channelsCó thể đã có người làm @psiddh đã nhận 3 ngày trước. Đang mởmodule: qnn module: quantization partner: qualcomm
pytorch/executorch#23160 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của pytorch/executorch
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-2 ngày Mức phù hợp với người mới 70/100
-
FingerprintSplitter raises ZeroDivisionError when int(frac_train * len(dataset)) floors to zeroĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 7 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
lmstudio-ai/mlx-engine#376 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
pyiron/bagofholding#166 ·