Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Measure vllm-tt-plugin gateability on the row's P150 (qwen35 served, tokens emitted)

Abierto
#3,261 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo

Línea de trabajo

Start with .agents/oracles/vllm-tt-plugin.md and the oracle-registry checker. Install vLLM 0.26.0 and the plugin at pin 7250ddfaa in a tt-metal environment on the row's P150, then serve one registered qwen35 checkpoint and record tokens emitted. Done means the oracle record can be marked gateable and includes the measurement; GGUF-decode parity is out of scope.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Row: BACKEND-TENSTORRENT

Local-Issue: ISSUE-LOCAL-01M32FVCE8XBW2PDTE48ZD8TWN
Kind: bug

The vllm-tt-plugin oracle record (.agents/oracles/vllm-tt-plugin.md, pin 7250ddfaa, 2026-09-21) is filed with gateable = no: nothing from the plugin has executed in this repository. The oracle-registry checker requires an owing issue named on a gateable=no record; this is it. Owed measurement: install vLLM 0.26.0 (VLLM_TARGET_DEVICE=empty) plus the plugin at the pinned HEAD inside a tt-metal environment on the row's P150, serve one registered qwen35 checkpoint, and record tokens emitted. Establishing gateability also unlocks the primary-rank denominators the Tenstorrent rows want: on-device qwen35 graph correctness at bf16 (vs our GGUF-decoded bf16 arm, per-layer like the llama.cpp bisect) and the native ~50 tok/s qwen35 rate measured side by side on the same card. Source: vllm.ai blog 2026-09-07; supported architectures include TTQwen3_5ForConditionalGeneration. NOT established by this measurement: GGUF-decode parity (the plugin serves HF weights; llama.cpp stays the GGUF decode oracle).

Lenguaje dominante
C++
Estrellas
423
Forks
53
Merge medio
1 d 7 h
PR fusionados (30 d)
382

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de mudler/vllm.cpp

Todos los issues de mudler/vllm.cpp

Issues similares

Más issues de C++

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.