[Bug] DFlash 2 speculative decoding crashes and fails in Python due to missing ctx_other and non-autoregressive graph execution
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Lĩnh vực
- ai, ai-infra-agents, backend
Hướng nghiên cứu
Issue nằm trong llama_cpp/llama.py và _internals.py, nơi việc khởi tạo mô hình draft DFlash 2 thất bại do thiếu liên kết ctx_other. Hãy kiểm tra fork C++ tại z-lab/llama.cpp-fork (commit 5ecbe1ac) để hiểu pipeline block-diffusion native. Bản sửa lỗi có thể sẽ bao gồm việc expose ctx_other trong Python bindings và đảm bảo draft model subclass tích hợp với C++ speculative decoding flow. Hãy kiểm thử bằng các model Qwen3.8-27B được cung cấp để xác minh rằng draft acceptance rates được cải thiện.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Prerequisites
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
When loading the target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF (Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf) together with the DFlash 2 draft model incoai/Qwen3.8-27B-DFlash2-GGUF (Qwen3.8-27B-DFlash2-Q4_K_M.gguf), llama-cpp-python should link the draft context to the target context via cparams.ctx_other = target_context and execute the native C++ block-diffusion speculative decoding pipeline, producing speculative speedups.
Current Behavior
- Calling
draft_llm = Llama(model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf")fails in GGML with:
dflash requires ctx_other to be set->ValueError: Failed to create llama_context
because DFlash 2 sidecar GGUF files do not contain their ownlm_head/output.weighttensors and requirectx_otherto be set during context initialization. - Setting
draft_llm.model = draft_model_ptrfails with:
AttributeError: property 'model' of 'Llama' object has no setter. - Calling
LlamaDraftModel(draft_llm, num_pred_tokens=5)fails with:
TypeError: LlamaDraftModel() takes no argumentsbecauseLlamaDraftModelis an abstract base class. - When subclassing
LlamaDraftModeland returning a Pythonlist, it crashes insidellama.pywith:
AttributeError: 'list' object has no attribute 'astype'. - When returning a
numpy.ndarraywithdtype=np.intc, the Python loop callsdraft_llm.sample(). This bypasses the C++ DFlash 2 pipeline (target hidden layer extractionllama_get_embeddings_layer_inp, encoder passllama_encode, and candidate lattice selectionbuild_post_sampling), resulting in a 0% draft acceptance rate and dropping generation speed to baseline autoregressive throughput (~34.27 tok/s).
Environment and Context
- Hardware: NVIDIA RTX PRO 6000 Blackwell Server Edition (
sm_120), 24+ GB VRAM - Environment: Hugging Face ZeroGPU Space
- Operating System: Debian GNU/Linux 13 (trixie) / Linux 6.12.94-123.192.amzn2023.x86_64
- glibc Version: Debian GLIBC 2.41-12
- Python Version: 3.12.12
- NVIDIA Driver: 580.159.03 (CUDA 13.0)
- llama-cpp-python: v0.3.35 compiled against submodule
vendor/llama.cppusing:- Repository: https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4
- Branch:
dflash2 - Commit Hash:
5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 - Upstream Pull Request: https://github.com/ggml-org/llama.cpp/pull/27342/changes
Steps to Reproduce
- Build
llama-cpp-pythonwithvendor/llama.cppchecked out atz-lab/llama.cpp-forkcommit5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4. - Download target model
utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUFand draft modelincoai/Qwen3.8-27B-DFlash2-GGUF. - Attempt to initialize the draft model in Python:
import llama_cpp
from llama_cpp import Llama
# 1. Target model initialization succeeds:
target_llm = Llama(
model_path="Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf",
n_gpu_layers=-1,
n_ctx=8192,
)
# 2. Draft model initialization fails here:
draft_llm = Llama(
model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf",
n_gpu_layers=-1,
n_ctx=8192,
)
Failure Logs
Context creation failure:
File "app.py", line 58, in get_or_load_model
draft_llm = Llama(model_path=DRAFT_MODEL_PATH, ...)
File ".../llama_cpp/llama.py", line 415, in __init__
internals.LlamaContext(
File ".../llama_cpp/_internals.py", line 266, in __init__
raise ValueError("Failed to create llama_context")
ValueError: Failed to create llama_context
Attribute setter failure:
File "app.py", line 94, in get_or_load_dflash2_model
draft_llm.model = draft_model_ptr
AttributeError: property 'model' of 'Llama' object has no setter
LlamaDraftModel constructor failure:
File "app.py", line 108, in get_or_load_dflash2_model
target_llm.draft_model = LlamaDraftModel(draft_model=draft_llm, num_pred_tokens=5)
TypeError: LlamaDraftModel() takes no arguments
List vs NumPy array failure:
File ".../llama_cpp/llama.py", line 1022, in generate
draft_tokens.astype(int)[
AttributeError: 'list' object has no attribute 'astype'
Additional Notes & Disclaimers
- Goal: I am trying to run DFlash 2 (
Qwen3.8-27B-DFlash2-Q4_K_M.gguf) on a Hugging Face ZeroGPU Space withllama-cpp-python. - Specific Fork: The C++ code is from https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 (PR https://github.com/ggml-org/llama.cpp/pull/27342/changes).
- Disclaimer on other speculative methods: The mention of EAGLE was suggested by an AI assistant during our discussion; I cannot personally confirm EAGLE's implementation details. I also do not know with certainty about built-in MTP, DFlash v1, or DSpark in
llama-cpp-python, though DFlash v1 and DSpark may already be present in upstreamllama.cpp. - AI Disclosure: This bug report was formatted with the help of an AI assistant at my request during a live debugging session. All steps, code snippets, stack traces, and reproduction logs come directly from my testing on Hugging Face ZeroGPU.
- Ngôn ngữ chính
- Python
- Star
- 10.6k
- Fork
- 1.5k
- Merge trung bình
- 3 giờ 57 phút
- Pull request đã merge (30 ngày)
- 4
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsCó thể đã có người làm @Belal0066 đã nhận 15 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
abetlen/llama-cpp-python#2371 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Có thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2352 · 1 bình luận · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Docs: consolidate build-from-source and GPU backend guideCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
abetlen/llama-cpp-python#2314 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2211 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyCó thể đã có người làm @Anai-Guo đã nhận 32 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2210 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của abetlen/llama-cpp-python
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 82/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 64/100
agrc/palletjack#208 ·
-
status/needs-triage type/bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
PKU-YuanGroup/OpenAI4S#218 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Test LeakageCó thể đã có người làm @garland3 đã nhận hôm nay. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
sandialabs/atlas-ui-3#1030 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
aws-samples/sample-ai-persona#151 ·
Maintainer thường phản hồi trong vòng 1 ngày