Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

perf(BACKEND-GATE-ROCM-SGLANG): close the Qwen3-4B Strix c4 gap

オープン
#3,076 コメント 14 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
リファクタリング
明瞭さ
説明が足りない
活発さ
活発
技術スタック
cpp

調査の方向性

まず .agents/specs/strix-qwen3-4b-c4-performance.md と #3053 の完全な記録を読み、その後、アイドル状態の Strix ホストで固定された Qwen3-4B BF16 c1/c4 の測定値を再現してください。引用されている rocm_embedding.hip、rocm.cpp、qwen3.cpp の箇所を評価する前に、同じツールを使って HIP/API/kernel の完全なトレースを収集してください。個別に測定した1つのレバーが必要な regression と gates を通過し、対応するウォームアップ済み c4 スループット、c1、レイテンシ、メモリ、正確性が報告されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Row: BACKEND-GATE-ROCM-SGLANG

Owner: current Strix campaign operator. Parent: #3053. Spec: .agents/specs/strix-qwen3-4b-c4-performance.md (commit before implementation).

User approved measurement-first work to bring vllm.cpp at least level with production vLLM at concurrency four on Strix. Two diagnostic qualification corpora gave medians 61.5947184118 versus 79.5022690204 output tokens/s. Correctness failed, so these are NOT_ACCEPTED diagnostic observations, not a benchmark or accepted ratio. Full record: #3053 issuecomment-5586767582.

Keep Qwen3-4B BF16 revision 1cfa9a7208912126459214e8b04321603b3df60c, six raw prompts, 128 greedy tokens, c1/c4, pinned engine baseline6e3cbfb940be89e28d1d71c264fd8c3a4e44afeb and production vLLMe126687a9a828d513c01a07cd69f025f27d63280. Obtain complete same-tool HIP/API/kernel traces before attribution. Trace06 is incomplete and cannot satisfy parity. Source candidates: rocm_embedding.hip127-142 per-step allocation/synchronization/copy/free; rocm.cpp91-97 and qwen3.cpp1184-1186 reject dense graph replay; paged attention BF16/head128/GQA4 does not use the fused-GQA arm. These are unmeasured hypotheses.

Coordinate existing graph #332/PR#2777 and async mirror PR#2779 before duplicating work. #3043 owns the separate 27B workload. Implement only a measured, individually specified lever, preserving invalid-input errors and production reachability. Require smallest red-before public-entry regression, focused and full gates, independent static and mutation review, operator verification, and a same-binary A/B on an idle leased host. Correctness requires the declared exact or separately ratified distributional gate. No eager denominator, candidate-driven tolerance, pin substitution, or ceiling claim. Final acceptance requires matched warmed repeated throughput at c4 >= vLLM, with c1, latency and memory obligations reported; otherwise keep the gap open and name the next traceable hypothesis.

主要言語
C++
スター
423
フォーク
53
平均マージ
1日 4時間
マージ済み PR(30日)
366

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

mudler/vllm.cpp のほかの issue

mudler/vllm.cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。