Require distinct Q and K projections in native HRX Qwen fusion
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 72/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- cpp
- Área
- ai-infra-agents, backend
Línea de trabajo
Empieza en ggml/src/ggml-hrx/dispatch_registration/qwen/dispatch-qwen-attention-postprocess.cpp, en match_qwen_attention_postprocess, alrededor de la línea 623; después revisa las pruebas del key-first scheduler en build-hrx-prefill-repair/evidence/key-first-scheduler-red.{json,log}. Verifica que el matcher requiera nodos de proyección de query y key distintos, manteniendo los defaults existentes y la fusión completa; luego vuelve a ejecutar la regresión de CPU con BF16 Q16/KV8 y un recuento de tokens de 1 para confirmar que no queda ningún RMSNorm genérico.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Row: BACKEND-ROCM
The pinned AMD llama.cpp Qwen attention matcher accepts a K projection as both the query and key chain. When a complete graph visits K before Q, it emits one partial Qwen postprocess and leaves the real query to generic RMSNorm and RoPE. match_qwen_attention_postprocess in ggml/src/ggml-hrx/dispatch_registration/qwen/dispatch-qwen-attention-postprocess.cpp:623 does not require distinct query and key projection nodes.
The #3083 implementer owns the in-flow scheduling correction. The reservation helper already rejects this ambiguity, but the actual matcher must enforce the same complete-region invariant. Preserve the original matching and traversal defaults, numerical tolerances, and the complete Qwen fusion.
The CPU regression enters the public dispatch scheduler with BF16 Q16/KV8, token count1, and K-before-Q graph order. It fails at the assertion that a complete Qwen postprocess leaves no generic RMSNorm: exit -6, test ELF 1fe0b0bbf1a22a84fb652816ad943e678e415d167da8bfb9383c3db9afb613a8, unchanged backend fb2301563318df2253bc681272247a2a5587a42921a9b376b89fa103de70b2a9. Evidence: build-hrx-prefill-repair/evidence/key-first-scheduler-red.{json,log} in the #3083 worktree.
The #3083 spec will record this narrow matcher obligation before its implementation. A separate sequential GPU regression currently rejects NaN in the first decode K cache. This issue does not yet attribute that numerical failure to this matcher ambiguity; the operator will rerun it after the scoped correction.
- Lenguaje dominante
- C++
- Estrellas
- 423
- Forks
- 53
- Merge medio
- 1 d 7 h
- PR fusionados (30 d)
- 380
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de mudler/vllm.cpp
-
[Windows] full build fails in tools/bench/conv1d_scaling_probe.cpp (POSIX-only sys/resource.h)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
Todos los issues de mudler/vllm.cpp
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
sandialabs/seacas#945 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
ROCm/FastFlowLM#757 ·
Los mantenedores suelen responder en 1 día
-
WaterHeaterManagement: tank_percent feature reports wrong feature id (FeatureMap corruption)Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
espressif/esp-matter#1867 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
mltframework/shotcut#1920 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día