Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

perf(BACKEND-GATE-ROCM-SGLANG): close the Qwen3-4B Strix c4 gap

Aperta
#3,076 14 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Refactoring
Chiarezza
Da chiarire
Stato di attività
Attiva
Stack tecnologico
cpp

Direzione di ricerca

Leggi prima .agents/specs/strix-qwen3-4b-c4-performance.md e il record completo in #3053, quindi riproduci le misurazioni fissate Qwen3-4B BF16 c1/c4 su un host Strix inattivo. Raccogli trace completi HIP/API/kernel con lo stesso strumento prima di valutare le posizioni citate in rocm_embedding.hip, rocm.cpp e qwen3.cpp. Il lavoro è completato quando una leva misurata singolarmente supera la regressione e i gate richiesti, e vengono riportati il throughput c4 a caldo corrispondente, c1, latenza, memoria e correttezza.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Row: BACKEND-GATE-ROCM-SGLANG

Owner: current Strix campaign operator. Parent: #3053. Spec: .agents/specs/strix-qwen3-4b-c4-performance.md (commit before implementation).

User approved measurement-first work to bring vllm.cpp at least level with production vLLM at concurrency four on Strix. Two diagnostic qualification corpora gave medians 61.5947184118 versus 79.5022690204 output tokens/s. Correctness failed, so these are NOT_ACCEPTED diagnostic observations, not a benchmark or accepted ratio. Full record: #3053 issuecomment-5586767582.

Keep Qwen3-4B BF16 revision 1cfa9a7208912126459214e8b04321603b3df60c, six raw prompts, 128 greedy tokens, c1/c4, pinned engine baseline6e3cbfb940be89e28d1d71c264fd8c3a4e44afeb and production vLLMe126687a9a828d513c01a07cd69f025f27d63280. Obtain complete same-tool HIP/API/kernel traces before attribution. Trace06 is incomplete and cannot satisfy parity. Source candidates: rocm_embedding.hip127-142 per-step allocation/synchronization/copy/free; rocm.cpp91-97 and qwen3.cpp1184-1186 reject dense graph replay; paged attention BF16/head128/GQA4 does not use the fused-GQA arm. These are unmeasured hypotheses.

Coordinate existing graph #332/PR#2777 and async mirror PR#2779 before duplicating work. #3043 owns the separate 27B workload. Implement only a measured, individually specified lever, preserving invalid-input errors and production reachability. Require smallest red-before public-entry regression, focused and full gates, independent static and mutation review, operator verification, and a same-binary A/B on an idle leased host. Correctness requires the declared exact or separately ratified distributional gate. No eager denominator, candidate-driven tolerance, pin substitution, or ceiling claim. Final acceptance requires matched warmed repeated throughput at c4 >= vLLM, with c1, latency and memory obligations reported; otherwise keep the gap open and name the next traceable hypothesis.

Lingua principale
C++
Stelle
440
Fork
55
Merge medio
1g 4h
PR unite (30g)
366

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di mudler/vllm.cpp

Tutte le issue di mudler/vllm.cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.