Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug] DFlash 2 speculative decoding crashes and fails in Python due to missing ctx_other and non-autoregressive graph execution

Đang mở
#2,361 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
25/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
c, python
Lĩnh vực
ai, ai-infra-agents, backend

Hướng nghiên cứu

Issue nằm trong llama_cpp/llama.py và _internals.py, nơi việc khởi tạo mô hình draft DFlash 2 thất bại do thiếu liên kết ctx_other. Hãy kiểm tra fork C++ tại z-lab/llama.cpp-fork (commit 5ecbe1ac) để hiểu pipeline block-diffusion native. Bản sửa lỗi có thể sẽ bao gồm việc expose ctx_other trong Python bindings và đảm bảo draft model subclass tích hợp với C++ speculative decoding flow. Hãy kiểm thử bằng các model Qwen3.8-27B được cung cấp để xác minh rằng draft acceptance rates được cải thiện.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Prerequisites

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Expected Behavior

When loading the target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF (Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf) together with the DFlash 2 draft model incoai/Qwen3.8-27B-DFlash2-GGUF (Qwen3.8-27B-DFlash2-Q4_K_M.gguf), llama-cpp-python should link the draft context to the target context via cparams.ctx_other = target_context and execute the native C++ block-diffusion speculative decoding pipeline, producing speculative speedups.

Current Behavior

  1. Calling draft_llm = Llama(model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf") fails in GGML with:
    dflash requires ctx_other to be set -> ValueError: Failed to create llama_context
    because DFlash 2 sidecar GGUF files do not contain their own lm_head / output.weight tensors and require ctx_other to be set during context initialization.
  2. Setting draft_llm.model = draft_model_ptr fails with:
    AttributeError: property 'model' of 'Llama' object has no setter.
  3. Calling LlamaDraftModel(draft_llm, num_pred_tokens=5) fails with:
    TypeError: LlamaDraftModel() takes no arguments because LlamaDraftModel is an abstract base class.
  4. When subclassing LlamaDraftModel and returning a Python list, it crashes inside llama.py with:
    AttributeError: 'list' object has no attribute 'astype'.
  5. When returning a numpy.ndarray with dtype=np.intc, the Python loop calls draft_llm.sample(). This bypasses the C++ DFlash 2 pipeline (target hidden layer extraction llama_get_embeddings_layer_inp, encoder pass llama_encode, and candidate lattice selection build_post_sampling), resulting in a 0% draft acceptance rate and dropping generation speed to baseline autoregressive throughput (~34.27 tok/s).

Environment and Context

Steps to Reproduce

  1. Build llama-cpp-python with vendor/llama.cpp checked out at z-lab/llama.cpp-fork commit 5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4.
  2. Download target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF and draft model incoai/Qwen3.8-27B-DFlash2-GGUF.
  3. Attempt to initialize the draft model in Python:
import llama_cpp
from llama_cpp import Llama

# 1. Target model initialization succeeds:
target_llm = Llama(
    model_path="Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf",
    n_gpu_layers=-1,
    n_ctx=8192,
)

# 2. Draft model initialization fails here:
draft_llm = Llama(
    model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf",
    n_gpu_layers=-1,
    n_ctx=8192,
)

Failure Logs

Context creation failure:

  File "app.py", line 58, in get_or_load_model
    draft_llm = Llama(model_path=DRAFT_MODEL_PATH, ...)
  File ".../llama_cpp/llama.py", line 415, in __init__
    internals.LlamaContext(
  File ".../llama_cpp/_internals.py", line 266, in __init__
    raise ValueError("Failed to create llama_context")
ValueError: Failed to create llama_context

Attribute setter failure:

  File "app.py", line 94, in get_or_load_dflash2_model
    draft_llm.model = draft_model_ptr
AttributeError: property 'model' of 'Llama' object has no setter

LlamaDraftModel constructor failure:

  File "app.py", line 108, in get_or_load_dflash2_model
    target_llm.draft_model = LlamaDraftModel(draft_model=draft_llm, num_pred_tokens=5)
TypeError: LlamaDraftModel() takes no arguments

List vs NumPy array failure:

  File ".../llama_cpp/llama.py", line 1022, in generate
    draft_tokens.astype(int)[
AttributeError: 'list' object has no attribute 'astype'

Additional Notes & Disclaimers

  • Goal: I am trying to run DFlash 2 (Qwen3.8-27B-DFlash2-Q4_K_M.gguf) on a Hugging Face ZeroGPU Space with llama-cpp-python.
  • Specific Fork: The C++ code is from https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 (PR https://github.com/ggml-org/llama.cpp/pull/27342/changes).
  • Disclaimer on other speculative methods: The mention of EAGLE was suggested by an AI assistant during our discussion; I cannot personally confirm EAGLE's implementation details. I also do not know with certainty about built-in MTP, DFlash v1, or DSpark in llama-cpp-python, though DFlash v1 and DSpark may already be present in upstream llama.cpp.
  • AI Disclosure: This bug report was formatted with the help of an AI assistant at my request during a live debugging session. All steps, code snippets, stack traces, and reproduction logs come directly from my testing on Hugging Face ZeroGPU.
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.5k
Merge trung bình
3 giờ 57 phút
Pull request đã merge (30 ngày)
4

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của abetlen/llama-cpp-python

Tất cả issue của abetlen/llama-cpp-python

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.