[Bug] DFlash 2 speculative decoding crashes and fails in Python due to missing ctx_other and non-autoregressive graph execution
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 25/100
- Tipo de issue
- Bug
- Clareza
- Claramente especificada
- Status de atividade
- Pouca atividade
- Domínio
- ai, ai-infra-agents, backend
Direção de pesquisa
A issue está em llama_cpp/llama.py e _internals.py, onde a inicialização do modelo draft DFlash 2 falha devido à ausência da vinculação com ctx_other. Examine o fork de C++ em z-lab/llama.cpp-fork (commit 5ecbe1ac) para entender o pipeline nativo de block-diffusion. A correção provavelmente envolve expor ctx_other nos bindings de Python e garantir que a subclasse do modelo draft se integre ao fluxo de speculative decoding em C++. Teste com os modelos Qwen3.8-27B fornecidos para verificar se as taxas de aceitação do draft melhoram.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Prerequisites
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
When loading the target model utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF (Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf) together with the DFlash 2 draft model incoai/Qwen3.8-27B-DFlash2-GGUF (Qwen3.8-27B-DFlash2-Q4_K_M.gguf), llama-cpp-python should link the draft context to the target context via cparams.ctx_other = target_context and execute the native C++ block-diffusion speculative decoding pipeline, producing speculative speedups.
Current Behavior
- Calling
draft_llm = Llama(model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf")fails in GGML with:
dflash requires ctx_other to be set->ValueError: Failed to create llama_context
because DFlash 2 sidecar GGUF files do not contain their ownlm_head/output.weighttensors and requirectx_otherto be set during context initialization. - Setting
draft_llm.model = draft_model_ptrfails with:
AttributeError: property 'model' of 'Llama' object has no setter. - Calling
LlamaDraftModel(draft_llm, num_pred_tokens=5)fails with:
TypeError: LlamaDraftModel() takes no argumentsbecauseLlamaDraftModelis an abstract base class. - When subclassing
LlamaDraftModeland returning a Pythonlist, it crashes insidellama.pywith:
AttributeError: 'list' object has no attribute 'astype'. - When returning a
numpy.ndarraywithdtype=np.intc, the Python loop callsdraft_llm.sample(). This bypasses the C++ DFlash 2 pipeline (target hidden layer extractionllama_get_embeddings_layer_inp, encoder passllama_encode, and candidate lattice selectionbuild_post_sampling), resulting in a 0% draft acceptance rate and dropping generation speed to baseline autoregressive throughput (~34.27 tok/s).
Environment and Context
- Hardware: NVIDIA RTX PRO 6000 Blackwell Server Edition (
sm_120), 24+ GB VRAM - Environment: Hugging Face ZeroGPU Space
- Operating System: Debian GNU/Linux 13 (trixie) / Linux 6.12.94-123.192.amzn2023.x86_64
- glibc Version: Debian GLIBC 2.41-12
- Python Version: 3.12.12
- NVIDIA Driver: 580.159.03 (CUDA 13.0)
- llama-cpp-python: v0.3.35 compiled against submodule
vendor/llama.cppusing:- Repository: https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4
- Branch:
dflash2 - Commit Hash:
5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 - Upstream Pull Request: https://github.com/ggml-org/llama.cpp/pull/27342/changes
Steps to Reproduce
- Build
llama-cpp-pythonwithvendor/llama.cppchecked out atz-lab/llama.cpp-forkcommit5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4. - Download target model
utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUFand draft modelincoai/Qwen3.8-27B-DFlash2-GGUF. - Attempt to initialize the draft model in Python:
import llama_cpp
from llama_cpp import Llama
# 1. Target model initialization succeeds:
target_llm = Llama(
model_path="Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf",
n_gpu_layers=-1,
n_ctx=8192,
)
# 2. Draft model initialization fails here:
draft_llm = Llama(
model_path="Qwen3.8-27B-DFlash2-Q4_K_M.gguf",
n_gpu_layers=-1,
n_ctx=8192,
)
Failure Logs
Context creation failure:
File "app.py", line 58, in get_or_load_model
draft_llm = Llama(model_path=DRAFT_MODEL_PATH, ...)
File ".../llama_cpp/llama.py", line 415, in __init__
internals.LlamaContext(
File ".../llama_cpp/_internals.py", line 266, in __init__
raise ValueError("Failed to create llama_context")
ValueError: Failed to create llama_context
Attribute setter failure:
File "app.py", line 94, in get_or_load_dflash2_model
draft_llm.model = draft_model_ptr
AttributeError: property 'model' of 'Llama' object has no setter
LlamaDraftModel constructor failure:
File "app.py", line 108, in get_or_load_dflash2_model
target_llm.draft_model = LlamaDraftModel(draft_model=draft_llm, num_pred_tokens=5)
TypeError: LlamaDraftModel() takes no arguments
List vs NumPy array failure:
File ".../llama_cpp/llama.py", line 1022, in generate
draft_tokens.astype(int)[
AttributeError: 'list' object has no attribute 'astype'
Additional Notes & Disclaimers
- Goal: I am trying to run DFlash 2 (
Qwen3.8-27B-DFlash2-Q4_K_M.gguf) on a Hugging Face ZeroGPU Space withllama-cpp-python. - Specific Fork: The C++ code is from https://github.com/z-lab/llama.cpp-fork/tree/5ecbe1ac17ec0484c5b44af0bd580cdc9c428ed4 (PR https://github.com/ggml-org/llama.cpp/pull/27342/changes).
- Disclaimer on other speculative methods: The mention of EAGLE was suggested by an AI assistant during our discussion; I cannot personally confirm EAGLE's implementation details. I also do not know with certainty about built-in MTP, DFlash v1, or DSpark in
llama-cpp-python, though DFlash v1 and DSpark may already be present in upstreamllama.cpp. - AI Disclosure: This bug report was formatted with the help of an AI assistant at my request during a live debugging session. All steps, code snippets, stack traces, and reproduction logs come directly from my testing on Hugging Face ZeroGPU.
- Linguagem predominante
- Python
- Estrelas
- 10.6k
- Forks
- 1.5k
- Merge médio
- 3h 57min
- PRs com merge (30d)
- 4
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Sem modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsTalvez já em andamento @Belal0066 assumiu há 16 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
abetlen/llama-cpp-python#2371 ·
Mantenedores costumam responder em até 1 dia
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Talvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2352 · 1 comentário · 2 reações ·
Mantenedores costumam responder em até 1 dia
-
Docs: consolidate build-from-source and GPU backend guideTalvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
abetlen/llama-cpp-python#2314 ·
Mantenedores costumam responder em até 1 dia
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_arrayTalvez livre de novo @lxcxjxhx assumiu há 93 dias e não há nenhum pull request aberto. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2211 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyTalvez já em andamento @Anai-Guo assumiu há 34 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
abetlen/llama-cpp-python#2210 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de abetlen/llama-cpp-python
Issues semelhantes
-
bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
Deepak3699/Ai_Mentor#244 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 82/100
btclib-org/btclib-node#1880 ·
Mantenedores costumam responder em até 1 dia
-
CONTRIBUTING.md: say how ticketless bug fixes and feature PRs are handledTalvez já em andamento @khuisman assumiu hoje. Abertav0.9.2
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 84/100
khuisman/mcp-gee-sweet#941 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
pyjanitor-devs/pyjanitor#1758 ·
Mantenedores costumam responder em até 1 dia
-
bug ready for review
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 86/100
odysseus-dev/odysseus#6641 ·
Mantenedores costumam responder em até 1 dia