Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Vulkan] LLM decode ~12% slower than XNNPACK on Adreno

未关闭
#22,968 13 条评论 2 个 reaction 已指派 2 人 在 GitHub 查看

维护者通常 1 天内回复

@digantdesai 已经在做这个了。

开始于 2026年9月21日。

评估

这个 Issue 还没有评估数据。

描述

module: vulkan triaged

Vulkan LLM decode is ~11-13% slower than XNNPACK on Adreno 840 (Galaxy S26 Ultra), while prefill is ~2x faster. Qwen3, llama_main, cold-device medians, arms interleaved.

Vulkan (4w) XNNPACK (8da4w)
0.6B decode 98.0 tok/s 109.0 tok/s
1.7B decode 41.3 tok/s 47.6 tok/s
0.6B prefill, 41-token prompt 1568 tok/s 818 tok/s

At 0.6B, decode reads 335 MB of weights per token. The q4gsw GEMV is ~74% of GPU time at ~44.5 GB/s, and the GPU is ~100% busy.

#22941 (fused QK+softmax on the SDPA decode path) is the one change that helped: +2.5-3.9% decode.

Tried, no gain:

  • 8da4w (linear_dq8ca_q4gsw): 22% slower than 4w
  • one command buffer per token: no change
  • fusing SwiGLU (sigmoid + 2 muls): 0.998x
  • tin GEMM at M=1 instead of the coop GEMV: 5.3x slower
  • int3 weights: would beat XNNPACK, but +40% perplexity even with HQQ+AWQ

Possible directions: cut the non-GEMV work (SDPA ~12%, rms_norm ~5%, no-op view copies ~5%), or a sub-4-bit linear kernel paired with a quantizer that holds accuracy.

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

主要语言
Python
星标
5k
派生
1.2k
平均合并
2 天 9 小时
30 天内合并 PR
555

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

pytorch/executorch 的其他 Issue

查看 pytorch/executorch 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。