Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Bug] DFlash 2 speculative decoding crashes and fails in Python due to missing ctx_other and non-autoregressive graph execution

Abierto
#2,361 1 comentario 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
25/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Tranquilo
Stack tecnológico
c, python

Línea de trabajo

El issue está en llama_cpp/llama.py y _internals.py, donde la inicialización del modelo draft DFlash 2 falla debido a la ausencia del enlace con ctx_other. Examina el fork de C++ en z-lab/llama.cpp-fork (commit 5ecbe1ac) para comprender el pipeline nativo de block-diffusion. Es probable que el fix implique exponer ctx_other en los bindings de Python y garantizar que la subclase del modelo draft se integre con el flujo de speculative decoding de C++. Prueba con los modelos Qwen3.8-27B proporcionados para verificar que mejoren las tasas de aceptación del draft.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Prerequisites

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Expected Behavior

When loading the target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF (Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf) together with the DFlash 2 draft model incoai/Qwen3.8-27B-DFlash2-GGUF (Qwen3.8-27B-DFlash2-Q4_K_M.gguf), llama-cpp-python should link the draft context to the target context via cparams.ctx_other = target_context and execute the native C++ block-diffusion speculative decoding pipeline, producing speculative speedups.

Current Behavior

  1. Calling draft_llm = Llama(model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf") fails in GGML with:
    dflash requires ctx_other to be set -> ValueError: Failed to create llama_context
    because DFlash 2 sidecar GGUF files do not contain their own lm_head / output.weight tensors and require ctx_other to be set during context initialization.
  2. Setting draft_llm.model = draft_model_ptr fails with:
    AttributeError: property 'model' of 'Llama' object has no setter.
  3. Calling LlamaDraftModel(draft_llm, num_pred_tokens=5) fails with:
    TypeError: LlamaDraftModel() takes no arguments because LlamaDraftModel is an abstract base class.
  4. When subclassing LlamaDraftModel and returning a Python list, it crashes inside llama.py with:
    AttributeError: 'list' object has no attribute 'astype'.
  5. When returning a numpy.ndarray with dtype=np.intc, the Python loop calls draft_llm.sample(). This bypasses the C++ DFlash 2 pipeline (target hidden layer extraction llama_get_embeddings_layer_inp, encoder pass llama_encode, and candidate lattice selection build_post_sampling), resulting in a 0% draft acceptance rate and dropping generation speed to baseline autoregressive throughput (~34.27 tok/s).

Environment and Context

Steps to Reproduce

  1. Build llama-cpp-python with vendor/llama.cpp checked out at z-lab/llama.cpp-fork commit 5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4.
  2. Download target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF and draft model incoai/Qwen3.8-27B-DFlash2-GGUF.
  3. Attempt to initialize the draft model in Python:
import llama_cpp
from llama_cpp import Llama

# 1. Target model initialization succeeds:
target_llm = Llama(
    model_path="Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf",
    n_gpu_layers=-1,
    n_ctx=8192,
)

# 2. Draft model initialization fails here:
draft_llm = Llama(
    model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf",
    n_gpu_layers=-1,
    n_ctx=8192,
)

Failure Logs

Context creation failure:

  File "app.py", line 58, in get_or_load_model
    draft_llm = Llama(model_path=DRAFT_MODEL_PATH, ...)
  File ".../llama_cpp/llama.py", line 415, in __init__
    internals.LlamaContext(
  File ".../llama_cpp/_internals.py", line 266, in __init__
    raise ValueError("Failed to create llama_context")
ValueError: Failed to create llama_context

Attribute setter failure:

  File "app.py", line 94, in get_or_load_dflash2_model
    draft_llm.model = draft_model_ptr
AttributeError: property 'model' of 'Llama' object has no setter

LlamaDraftModel constructor failure:

  File "app.py", line 108, in get_or_load_dflash2_model
    target_llm.draft_model = LlamaDraftModel(draft_model=draft_llm, num_pred_tokens=5)
TypeError: LlamaDraftModel() takes no arguments

List vs NumPy array failure:

  File ".../llama_cpp/llama.py", line 1022, in generate
    draft_tokens.astype(int)[
AttributeError: 'list' object has no attribute 'astype'

Additional Notes & Disclaimers

  • Goal: I am trying to run DFlash 2 (Qwen3.8-27B-DFlash2-Q4_K_M.gguf) on a Hugging Face ZeroGPU Space with llama-cpp-python.
  • Specific Fork: The C++ code is from https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 (PR https://github.com/ggml-org/llama.cpp/pull/27342/changes).
  • Disclaimer on other speculative methods: The mention of EAGLE was suggested by an AI assistant during our discussion; I cannot personally confirm EAGLE's implementation details. I also do not know with certainty about built-in MTP, DFlash v1, or DSpark in llama-cpp-python, though DFlash v1 and DSpark may already be present in upstream llama.cpp.
  • AI Disclosure: This bug report was formatted with the help of an AI assistant at my request during a live debugging session. All steps, code snippets, stack traces, and reproduction logs come directly from my testing on Hugging Face ZeroGPU.
Lenguaje dominante
Python
Estrellas
10.6k
Forks
1.5k
Merge medio
3 h 57 min
PR fusionados (30 d)
4

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de abetlen/llama-cpp-python

Todos los issues de abetlen/llama-cpp-python

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.