Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[ET-VK] is_coopmat_eligible() disables cooperative matrix on all integrated GPUs

オープン
#22,665 コメント 6 件 リアクション 1 件 担当者 1 名 GitHub で見る

@SS-JIA がすでに取り組んでいます。

2026年9月10日 から。

評価

この issue はまだ評価されていません。

説明

enhancement module: vulkan needs reproduction
🐛 Describe the bug

is_coopmat_eligible() in backends/vulkan/runtime/graph/ops/impl/GemmCoopmat.h disables the cooperative-matrix GEMM path on every integrated GPU, regardless of whether the device advertises VK_KHR_cooperative_matrix:

inline bool is_coopmat_eligible(
    ComputeGraph& graph, const ValueRef out, int64_t M, int64_t N, int64_t K) {
  if (graph.dim_of(out) > 2) {
    return false;
  }
  const auto* adapter = graph.context()->adapter_ptr();
  return adapter->supports_cooperative_matrix() &&
      adapter->subgroup_size() == 64 && !adapter->is_integrated_gpu() &&
      graph.storage_type_of(out) == utils::kBuffer && M % kCoopmatTileM == 0 &&
      N % kCoopmatTileN == 0 && K % kCoopmatTileK == 0;
}

On an AMD Strix Halo (Radeon 8060S, RADV, GFX1151) every other condition is satisfied — the device reports cooperative matrix support and a subgroup size of 64 — but !adapter->is_integrated_gpu() rejects it. For reference, llama.cpp's Vulkan backend enumerates the same device as

Radeon 8060S Graphics (RADV GFX1151) (radv) | uma: 1 | fp16: 1 | bf16: 0 | warp size: 64 | matrix cores: KHR_coopmat

and uses coopmat on it.

A second condition, dim_of(out) > 2, separately excludes batch-1 transformer activations, which are (1, M, N). A tensor whose leading dimensions are all 1 has the same buffer layout as the 2-D tensor it wraps.

Is the integrated-GPU exclusion deliberate — guarding a known correctness or performance problem on APUs — or is it a conservative default? Nothing in the source or in #19009 says which, which is why this is a question rather than a PR.

What I measured

I removed both conditions locally (dropped !adapter->is_integrated_gpu(), and relaxed the rank check to permit leading dimensions equal to 1) and ran the V-JEPA 2 ViT-L encoder, 512 tokens, fp32, buffer storage. Identical .pte on both runtimes, so only the eligibility check differs:

graph stock coopmat enabled
heavily partitioned (73 delegate calls) 487.8 ms 427.2 ms (−12.4%)
single delegate call 141.3 ms 140.3 ms (no change)

Accuracy was neutral, measured over five different inputs against eager PyTorch:

stock coopmat enabled
relative L2 2.845% – 4.832% 2.807% – 4.757%
worst-token cosine 0.86150 – 0.98976 0.87001 – 0.99032

So on this device the exclusion does not appear to be guarding a correctness problem, and the gain is real but workload-dependent — it showed up on a badly partitioned graph and not on a well-partitioned one, which is consistent with GEMM being a larger share of the former.

Context for the motivation: this encoder runs at roughly 2.4 TFLOP/s on this part, about 16% of its fp32 peak, with the whole graph in a single delegate call — so the time is in the shaders rather than at the partition boundary.

I'm happy to send a PR if the exclusion turns out to be conservative rather than intentional. I did not open one directly because I can only test a single integrated GPU, and a device-class gate is exactly the kind of thing that wants broader validation than I can give it.

Versions
  • ExecuTorch b20f16a70b8d9d9953c7ab15c05a1f5584cc36ef (1.4.0a0); the code is unchanged on main at the time of filing
  • torch 2.12.1+cpu, Ubuntu 24.04
  • AMD Strix Halo, Radeon 8060S (RADV GFX1151), Mesa 25.2.8
  • EXECUTORCH_BUILD_VULKAN=ON, buffer storage, fp32

cc @SS-JIA @manuelcandales @digantdesai @cbilgin

主要言語
Python
スター
5k
フォーク
1.2k
平均マージ
2日 13時間
マージ済み PR(30日)
595

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

pytorch/executorch のほかの issue

pytorch/executorch の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。