Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

[Bug] DFlash 2 speculative decoding crashes and fails in Python due to missing ctx_other and non-autoregressive graph execution

未關閉
#2,361 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
25/100
Issue 類型
缺陷
描述清晰度
描述清楚
活躍度
冷清
技術堆疊
c, python

研究方向

此 issue 位於 llama_cpp/llama.py 和 _internals.py 中,DFlash 2 draft model 的初始化因缺少 ctx_other 連結而失敗。檢查 z-lab/llama.cpp-fork 中的 C++ fork(commit 5ecbe1ac),以了解 native block-diffusion pipeline。修正可能涉及在 Python bindings 中公開 ctx_other,並確保 draft model subclass 與 C++ speculative decoding flow 整合。使用提供的 Qwen3.8-27B models 進行測試,以驗證 draft acceptance rates 得到提升。

由索引模型根據 Issue 內容生成。

描述

Prerequisites

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Expected Behavior

When loading the target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF (Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf) together with the DFlash 2 draft model incoai/Qwen3.8-27B-DFlash2-GGUF (Qwen3.8-27B-DFlash2-Q4_K_M.gguf), llama-cpp-python should link the draft context to the target context via cparams.ctx_other = target_context and execute the native C++ block-diffusion speculative decoding pipeline, producing speculative speedups.

Current Behavior

  1. Calling draft_llm = Llama(model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf") fails in GGML with:
    dflash requires ctx_other to be set -> ValueError: Failed to create llama_context
    because DFlash 2 sidecar GGUF files do not contain their own lm_head / output.weight tensors and require ctx_other to be set during context initialization.
  2. Setting draft_llm.model = draft_model_ptr fails with:
    AttributeError: property 'model' of 'Llama' object has no setter.
  3. Calling LlamaDraftModel(draft_llm, num_pred_tokens=5) fails with:
    TypeError: LlamaDraftModel() takes no arguments because LlamaDraftModel is an abstract base class.
  4. When subclassing LlamaDraftModel and returning a Python list, it crashes inside llama.py with:
    AttributeError: 'list' object has no attribute 'astype'.
  5. When returning a numpy.ndarray with dtype=np.intc, the Python loop calls draft_llm.sample(). This bypasses the C++ DFlash 2 pipeline (target hidden layer extraction llama_get_embeddings_layer_inp, encoder pass llama_encode, and candidate lattice selection build_post_sampling), resulting in a 0% draft acceptance rate and dropping generation speed to baseline autoregressive throughput (~34.27 tok/s).

Environment and Context

Steps to Reproduce

  1. Build llama-cpp-python with vendor/llama.cpp checked out at z-lab/llama.cpp-fork commit 5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4.
  2. Download target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF and draft model incoai/Qwen3.8-27B-DFlash2-GGUF.
  3. Attempt to initialize the draft model in Python:
import llama_cpp
from llama_cpp import Llama

# 1. Target model initialization succeeds:
target_llm = Llama(
    model_path="Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf",
    n_gpu_layers=-1,
    n_ctx=8192,
)

# 2. Draft model initialization fails here:
draft_llm = Llama(
    model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf",
    n_gpu_layers=-1,
    n_ctx=8192,
)

Failure Logs

Context creation failure:

  File "app.py", line 58, in get_or_load_model
    draft_llm = Llama(model_path=DRAFT_MODEL_PATH, ...)
  File ".../llama_cpp/llama.py", line 415, in __init__
    internals.LlamaContext(
  File ".../llama_cpp/_internals.py", line 266, in __init__
    raise ValueError("Failed to create llama_context")
ValueError: Failed to create llama_context

Attribute setter failure:

  File "app.py", line 94, in get_or_load_dflash2_model
    draft_llm.model = draft_model_ptr
AttributeError: property 'model' of 'Llama' object has no setter

LlamaDraftModel constructor failure:

  File "app.py", line 108, in get_or_load_dflash2_model
    target_llm.draft_model = LlamaDraftModel(draft_model=draft_llm, num_pred_tokens=5)
TypeError: LlamaDraftModel() takes no arguments

List vs NumPy array failure:

  File ".../llama_cpp/llama.py", line 1022, in generate
    draft_tokens.astype(int)[
AttributeError: 'list' object has no attribute 'astype'

Additional Notes & Disclaimers

  • Goal: I am trying to run DFlash 2 (Qwen3.8-27B-DFlash2-Q4_K_M.gguf) on a Hugging Face ZeroGPU Space with llama-cpp-python.
  • Specific Fork: The C++ code is from https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 (PR https://github.com/ggml-org/llama.cpp/pull/27342/changes).
  • Disclaimer on other speculative methods: The mention of EAGLE was suggested by an AI assistant during our discussion; I cannot personally confirm EAGLE's implementation details. I also do not know with certainty about built-in MTP, DFlash v1, or DSpark in llama-cpp-python, though DFlash v1 and DSpark may already be present in upstream llama.cpp.
  • AI Disclosure: This bug report was formatted with the help of an AI assistant at my request during a live debugging session. All steps, code snippets, stack traces, and reproduction logs come directly from my testing on Hugging Face ZeroGPU.
主要語言
Python
星號
10.6k
分支
1.5k
平均合併
3 小時 57 分鐘
30 天內合併 PR
4

環境準備

  • 沒有 Dockerfile 或 Docker Compose 檔案
  • 沒有 Pull Request 範本
  • 閱讀貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

abetlen/llama-cpp-python 的其他 Issue

查看 abetlen/llama-cpp-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。