Measure vllm-tt-plugin gateability on the row's P150 (qwen35 served, tokens emitted)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Área
- machine-learning
Línea de trabajo
Start with .agents/oracles/vllm-tt-plugin.md and the oracle-registry checker. Install vLLM 0.26.0 and the plugin at pin 7250ddfaa in a tt-metal environment on the row's P150, then serve one registered qwen35 checkpoint and record tokens emitted. Done means the oracle record can be marked gateable and includes the measurement; GGUF-decode parity is out of scope.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Row: BACKEND-TENSTORRENT
Local-Issue: ISSUE-LOCAL-01M32FVCE8XBW2PDTE48ZD8TWN
Kind: bug
The vllm-tt-plugin oracle record (.agents/oracles/vllm-tt-plugin.md, pin 7250ddfaa, 2026-09-21) is filed with gateable = no: nothing from the plugin has executed in this repository. The oracle-registry checker requires an owing issue named on a gateable=no record; this is it. Owed measurement: install vLLM 0.26.0 (VLLM_TARGET_DEVICE=empty) plus the plugin at the pinned HEAD inside a tt-metal environment on the row's P150, serve one registered qwen35 checkpoint, and record tokens emitted. Establishing gateability also unlocks the primary-rank denominators the Tenstorrent rows want: on-device qwen35 graph correctness at bf16 (vs our GGUF-decoded bf16 arm, per-layer like the llama.cpp bisect) and the native ~50 tok/s qwen35 rate measured side by side on the same card. Source: vllm.ai blog 2026-09-07; supported architectures include TTQwen3_5ForConditionalGeneration. NOT established by this measurement: GGUF-decode parity (the plugin serves HF weights; llama.cpp stays the GGUF decode oracle).
- Lenguaje dominante
- C++
- Estrellas
- 423
- Forks
- 53
- Merge medio
- 1 d 7 h
- PR fusionados (30 d)
- 382
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de mudler/vllm.cpp
-
[Windows] full build fails in tools/bench/conv1d_scaling_probe.cpp (POSIX-only sys/resource.h)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
Todos los issues de mudler/vllm.cpp
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
hyprwm/aquamarine#426 ·
Los mantenedores suelen responder en 1 día
-
Winget hash mismatch for 5.0.3.0Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
amnezia-vpn/amnezia-client#3222 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
valkey-io/valkey-search#1465 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
KhronosGroup/Vulkan-Tutorial#524 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
microsoft/onnxruntime-genai#2633 ·
Los mantenedores suelen responder en 1 día