Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Qualify HRX dense WMMA after Loom removes a live template provider

オープン
#3,081 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
cpp
領域
compilers

調査の方向性

#3080 の証拠から始め、変更していない loom-link と loom-compile を使って失敗を再現し、その後 loom/src/loom/transforms/symbol/template_selection.c:585 と inline-callables の挙動を調べます。適格性確認として test-hrx-ops と正確な dense WMMA recipe を使用します。完了の条件は、修正前に upstream-native regression が失敗し、live provider dependencies がインライン化後も維持され、正確な recipe がコンパイルでき、独立したレビューと native operation test に合格することです。

索引モデルが issue の本文から書いたものです。

説明

Row: BACKEND-ROCM

Native HRX qualification for #3080 fails because Loom removes a live template provider during dense WMMA compilation. This issue owns the bounded compiler repair and its qualification evidence.

At ROCm/hrx-system 6bcd5a4ff111fa5bf160ab9f4592ca8e7cc810b1 and AMD-Ecosystem/llama.cpp 6319038132ed12f968ea68f37753f705da830ea8, test-hrx-ops on an RX 7900 XTX fails at the first Q4_K dense WMMA case: input_size=2048, output_size=128, token_count=2, output_accumulation=0, output_unary_op=23, weight_format=4. The native backend detects gfx1100 successfully.

Unmodified loom-link plus loom-compile reproduces the same failure without a GPU. The first template selection succeeds. inline-callables then removes template definitions while scalar template calls to ggml_unary_f32_apply remain. The next selection reports template.call callee has no template provider contract at loom/src/loom/transforms/symbol/template_selection.c:585.

The owning operator retains the exact linked recipe, pass snapshots, source pins, commands, and failed GPU run under the #3080 evidence. The repair must add an upstream-native regression that fails before the change, preserve live provider dependencies through inlining, compile the exact dense recipe, and pass independent review. Root reruns the native operation test. Do not weaken a verifier, tolerance, or operation contract, and do not treat this failure as a gfx1100 hardware limitation.

Use a committed scoped spec before implementation. The #3080 operator owns the work and the eventual disposition of the upstream patch. Native performance remains unaccepted until its declared correctness gate passes.

主要言語
C++
スター
440
フォーク
55
平均マージ
1日 4時間
マージ済み PR(30日)
346

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

mudler/vllm.cpp のほかの issue

mudler/vllm.cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。