Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

KV-WARMUP-PROFILE: reconcile actual startup and generation config observations

Aperta
#3,061 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp, python

Direzione di ricerca

Leggi warmup-memory-profiling.md Gate 0, quindi esegui il red-before CPU reproducer sulle definizioni fissate. Traccia la transizione tra upstream_state_observer.py:86–92, vllm/config/compilation.py:37–48, vllm/platforms/rocm.py:894 e vllm/config/vllm.py:1427. Il lavoro è completato quando la strumentazione preserva il rifiuto rigoroso e le prove grezze ordinate, e supera i gate di mutation specificata, CPU/image, review e physical-control.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Row: KV-WARMUP-PROFILE

The private startup observer for #3046 refuses the first fresh physical control after the actual graph-manager repair #3059. The unchanged pinned vLLM model run succeeds, but the separate finalizer exits 1 with PROFILE3046_REFUSED startup/generation configuration differs: compilation.custom_ops. This is an instrument join failure, not an accepted memory measurement or a native model failure.

The operator ran source seal 39072f980e4028863c417eb78e83a9a11205f42cbeac14681fc55db05be8c1da after fresh review and independent CPU gates: 50 permanent tests, 8 upstream cases, 83 mutations, 31 paired controls, and 20 approved-image CPU methods. The native auto-KV control ran under /home/vikash/gpu.lock from 2026-09-08 05:50:50 to 05:58:00 UTC and exited 0. The finalizer exited 1. Source, model, and state-instrument bytes remain unchanged. The other eight prepared controls were not executed.

A read-only enumeration using the unchanged finalizer helpers finds two differences among all 30 overlapping fields:

  • compilation.custom_ops: startup records ["+sparse_attn_indexer", "none"]; generation records ["+sparse_attn_indexer", "none", "+sparse_attn_indexer"].
  • compilation.mode: startup comparison produces "VLLM_COMPILE"; generation records integer 3.

The active pin is e126687a9a828d513c01a07cd69f025f27d63280. vllm/config/compilation.py:37–48 defines CompilationMode as an IntEnum with VLLM_COMPILE=3. The accepted state observer's plain function takes integer subclasses before Enum at upstream_state_observer.py:86–92; the startup finalizer's comparable returns the enum name. The permanent fixtures did not expose this actual representation difference.

For the custom-op change, vllm/platforms/rocm.py:894 unconditionally appends +sparse_attn_indexer, called from VllmConfig.__post_init__ at vllm/config/vllm.py:1427. These are source leads. The actual intervening call and its object identity have not yet been established. Do not deduplicate, ignore this field, infer harmlessness, or accept changed dispatch without tracing the complete transition.

Owner: the gfx1100 campaign operator and a fresh private observer implementer, row KV-WARMUP-PROFILE. The committed warmup-memory-profiling.md Gate 0 requires exact source-bound startup/generation identity. Repair the instrument against that contract, with real pinned definitions and both actual encodings in the CPU reproducer. Preserve strict refusal for changed modes, custom-op selections, phases, workers, source identities, and incomplete observations. Keep raw ordered custom-op lists and all original failed evidence. No upstream/runtime/model/config/default changes or state-observer replacement are authorized by this issue.

Acceptance: red-before reproduction, smallest complete repair, effective mutations with paired detector-removal/restoration controls, complete CPU/image gates, fresh immutable-head review, operator rerun, and new physical controls. The original physical run remains refused. The owning integration closes the issue when the repaired instrumentation and evidence are accepted.

Evidence: profile-gap-ownership/observer-implementation/runtime-graph-manager-repair-3059/operator/startup-generation-config-differences.json, SHA256 854bc04bc4fd1c8272342b76f75062e30cc6b6f93cf840931fa41ad16d54672c. The corresponding gpu-runs/auto-enabled-1 contains the exact Docker vectors, actual launcher PID witness, native log, finalizer refusal, and all raw file hashes. Approved image: sha256:80aab4c182a1f3eeebe286173977e57fcaf10a049b41f475655b35d285de31dc. Model: pinned Qwen3.5-0.8B BF16, native production graphs, utilization 0.92, unchanged scheduling, eight generated tokens.

Lingua principale
C++
Stelle
423
Fork
53
Merge medio
1g 7h
PR unite (30g)
380

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di mudler/vllm.cpp

Tutte le issue di mudler/vllm.cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.