Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Measure vllm-tt-plugin gateability on the row's P150 (qwen35 served, tokens emitted)

Aperta
#3,261 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
35/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva

Direzione di ricerca

Start with .agents/oracles/vllm-tt-plugin.md and the oracle-registry checker. Install vLLM 0.26.0 and the plugin at pin 7250ddfaa in a tt-metal environment on the row's P150, then serve one registered qwen35 checkpoint and record tokens emitted. Done means the oracle record can be marked gateable and includes the measurement; GGUF-decode parity is out of scope.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Row: BACKEND-TENSTORRENT

Local-Issue: ISSUE-LOCAL-01M32FVCE8XBW2PDTE48ZD8TWN
Kind: bug

The vllm-tt-plugin oracle record (.agents/oracles/vllm-tt-plugin.md, pin 7250ddfaa, 2026-09-21) is filed with gateable = no: nothing from the plugin has executed in this repository. The oracle-registry checker requires an owing issue named on a gateable=no record; this is it. Owed measurement: install vLLM 0.26.0 (VLLM_TARGET_DEVICE=empty) plus the plugin at the pinned HEAD inside a tt-metal environment on the row's P150, serve one registered qwen35 checkpoint, and record tokens emitted. Establishing gateability also unlocks the primary-rank denominators the Tenstorrent rows want: on-device qwen35 graph correctness at bf16 (vs our GGUF-decoded bf16 arm, per-layer like the llama.cpp bisect) and the native ~50 tok/s qwen35 rate measured side by side on the same card. Source: vllm.ai blog 2026-09-07; supported architectures include TTQwen3_5ForConditionalGeneration. NOT established by this measurement: GGUF-decode parity (the plugin serves HF weights; llama.cpp stays the GGUF decode oracle).

Lingua principale
C++
Stelle
423
Fork
53
Merge medio
1g 7h
PR unite (30g)
380

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di mudler/vllm.cpp

Tutte le issue di mudler/vllm.cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.