[Vulkan] LLM decode ~12% slower than XNNPACK on Adreno
维护者通常 1 天内回复
@digantdesai 已经在做这个了。
开始于 2026年9月21日。
评估
这个 Issue 还没有评估数据。
描述
Vulkan LLM decode is ~11-13% slower than XNNPACK on Adreno 840 (Galaxy S26 Ultra), while prefill is ~2x faster. Qwen3, llama_main, cold-device medians, arms interleaved.
| Vulkan (4w) | XNNPACK (8da4w) | |
|---|---|---|
| 0.6B decode | 98.0 tok/s | 109.0 tok/s |
| 1.7B decode | 41.3 tok/s | 47.6 tok/s |
| 0.6B prefill, 41-token prompt | 1568 tok/s | 818 tok/s |
At 0.6B, decode reads 335 MB of weights per token. The q4gsw GEMV is ~74% of GPU time at ~44.5 GB/s, and the GPU is ~100% busy.
#22941 (fused QK+softmax on the SDPA decode path) is the one change that helped: +2.5-3.9% decode.
Tried, no gain:
- 8da4w (
linear_dq8ca_q4gsw): 22% slower than 4w - one command buffer per token: no change
- fusing SwiGLU (sigmoid + 2 muls): 0.998x
- tin GEMM at M=1 instead of the coop GEMV: 5.3x slower
- int3 weights: would beat XNNPACK, but +40% perplexity even with HQQ+AWQ
Possible directions: cut the non-GEMV work (SDPA ~12%, rms_norm ~5%, no-op view copies ~5%), or a sub-4-bit linear kernel paired with a quantizer that holds accuracy.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin
- 主要语言
- Python
- 星标
- 5k
- 派生
- 1.2k
- 平均合并
- 2 天 9 小时
- 30 天内合并 PR
- 555
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
pytorch/executorch 的其他 Issue
-
enhancement triaged
难度 2/5 半天 新手友好度 68/100
pytorch/executorch#21640 ·
维护者通常 1 天内回复
-
enhancement module: examples
难度 5/5 一周以上 新手友好度 20/100
pytorch/executorch#23164 · 7 条评论 · 1 个 reaction ·
维护者通常 1 天内回复
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channels可能已有人在做 @psiddh 于 2 天前认领。 未关闭module: qnn partner: qualcomm
pytorch/executorch#23160 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[cpu kernels] native_layer_norm: layer_norm_scalar returns NaN on large-mean rows; Half/BF16 at N>=256 slow after #23153可能已有人在做 @JakeStevens 于 2 天前认领。 未关闭module: kernels
pytorch/executorch#23159 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
module: vulkan
难度 3/5 1-2 天 新手友好度 66/100
pytorch/executorch#23158 ·
维护者通常 1 天内回复
查看 pytorch/executorch 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 76/100
PedestrianDynamics/pyFDS-Evac#199 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
521xueweihan/HelloGitHub#3790 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
sandialabs/atlas-ui-3#978 ·
维护者通常 1 天内回复
-
area: tests perceived difficulty: 2
难度 2/5 1-3 小时 新手友好度 72/100
Nitjsefnie-Harness-Commons/daedalus#1255 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 86/100
EleutherAI/lm-evaluation-harness#4256 ·
维护者通常 1 天内回复