Enable native Kev inference with the Vulkan backend
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 30/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Área
- ai-infra-agents, backend, desktop-dev, mobile-dev
Línea de trabajo
Start with the Kev example in examples/kev/ and the Vulkan export example in examples/vulkan/. Examine the Vulkan backend in backends/vulkan/ to identify missing operator support for GatedDeltaNet and dynamic shapes. Write a small GDN parity test as a first step, then extend the export and CMake linkage to allow kev_runner to use the Vulkan backend. Validate against the PyTorch reference implementation.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Enable the native Kev example on the Vulkan backend, extending the XNNPACK and MLX support added in #23023.
Kev runs prefill followed by pointer-head scoring, with an immutable prefix snapshot reused across question batches. The goal is to support this workflow through the existing C++ API on Vulkan-capable GPUs, without Python at inference.
Start with Kev's model/export code, the Vulkan export example, and the Vulkan backend. Qwen3.5 MoE's GatedDeltaNet implementation provides a reference for GDN math and state layouts; Vulkan will need its own lowering or shader support where coverage is missing.
Work to cover:
- Identify operator and dynamic-shape gaps in Kev's
prefillandscoregraphs. Add the required Vulkan support, including GDN, with focused backend regression tests. A small GDN parity test is a useful first step. - Add a Vulkan export option and CMake linkage so the existing
kev_runnerandkev_benchmarkrun the exported model throughModule. - Preserve
system_one, explicitprefill/evaluate, configurable token limits, variable question/option counts, and batching beyond eight questions. Repeated evaluations must leave the prefix unchanged. - Start with an unquantized FP32 path, retaining the fitted temperature and checkpoint metadata. Check logits, probabilities, and reused state against upstream Kev's PyTorch implementation and the existing FP32 path, with documented tolerances.
- Verify the backbone, including GDN, executes on Vulkan and report any CPU fallback. Document a tested GPU, precision requirements, export/build/run commands, and measurements using
kev_benchmark.
Keep the integration in the existing example and reuse the runner and benchmark. A Vulkan-capable GPU and the Vulkan SDK are needed for validation.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin @iseeyuan @lucylq @helunwencser @tarun292 @kimishpatel @jackzhxng
- Lenguaje dominante
- Python
- Estrellas
- 5k
- Forks
- 1.2k
- Merge medio
- 2 d 9 h
- PR fusionados (30 d)
- 555
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de pytorch/executorch
-
enhancement triaged
Dificultad 2/5 Medio día Aptitud para principiantes 68/100
pytorch/executorch#21640 ·
Los mantenedores suelen responder en 1 día
-
enhancement module: examples
Dificultad 5/5 Más de una semana Aptitud para principiantes 20/100
pytorch/executorch#23164 · 7 comentarios · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channelsPosiblemente ocupada @psiddh la tomó hace 2 días. Abiertomodule: qnn partner: qualcomm
pytorch/executorch#23160 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[cpu kernels] native_layer_norm: layer_norm_scalar returns NaN on large-mean rows; Half/BF16 at N>=256 slow after #23153Posiblemente ocupada @JakeStevens la tomó hace 2 días. Abiertomodule: kernels
pytorch/executorch#23159 · 2 comentarios · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
module: vulkan
Dificultad 3/5 1-2 días Aptitud para principiantes 66/100
pytorch/executorch#23158 ·
Los mantenedores suelen responder en 1 día
Todos los issues de pytorch/executorch
Issues similares
-
bug status/needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
prowler-cloud/prowler#12887 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
use-agent-os/agent-os#3484 ·
Los mantenedores suelen responder en 2 días
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
open-telemetry/opentelemetry-python-contrib#5113 · 2 comentarios · 2 reacciones ·
Los mantenedores suelen responder en 1 día
-
external
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
langchain-ai/docs#6255 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día