Multiple calls to create_chat_completion() fail with "llama_decode: failed to decode, ret = -1"
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 新手友好度
- 55/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 停滞
- 技术栈
- python
- 领域
- ai, backend-api-design
调研方向
该 issue 位于 llama_cpp.Llama 类的 create_chat_completion 方法中。查看该方法的源代码以及底层的 C++ bindings,了解 context 和 KV cache 在多次调用之间是如何管理的。用户发现调用 llm.reset() 或 llm._ctx.kv_cache_clear() 有效,因此修复可能涉及在一次 completion 后自动重置 state。检查 test suite 中是否有类似模式。“完成”的标准是:连续两次调用 create_chat_completion 都能成功,且无需手动 reset。
由索引模型根据 Issue 内容生成。
描述
Prerequisites
Please answer the following questions for yourself before submitting an issue.
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
I am running LiquidAI's LFM2.5-1.2B-Instruct model. Calling create_chat_completion() multiple times should not throw error.
Current Behavior
When trying to call create_chat_completion twice in a row, the model throws error "llama_decode: failed to decode, ret = -1"
Environment and Context
I am using llama-cpp-python v0.3.16. Python version is 3.13.9.
- Windows 11
Failure Information (for bugs)
On tracing back the issue, it looks like there needs to be a context and cache reset after each chat_completion call, which isn't happening yet.
Steps to Reproduce
from pathlib import Path
from llama_cpp import Llama
llm = Llama(
model_path=str(Path.home() / "AppData/Local/llama.cpp/LiquidAI_LFM2.5-1.2B-Instruct-GGUF_LFM2.5-1.2B-Instruct-Q4_K_M.gguf"),
n_ctx=1000
)
system_prompt = """
\nYou are a helpful assistant
"""
prompt = """
suggest me places to visit during winter season
"""
response = llm.create_chat_completion(
messages = [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)
# llm.reset() # Using this works
# llm._ctx.kv_cache_clear() # Using this works
response = llm.create_chat_completion(
messages = [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)
Failure Logs
init: the tokens of sequence 0 in the input batch have inconsistent sequence positions:
- the last position stored in the memory module of the context (i.e. the KV cache) for sequence 0 is X = 519
- the tokens for sequence 0 in the input batch have a starting position of Y = 29
it is required that the sequence positions remain consecutive: Y = X + 1
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
- 主要语言
- Python
- 星标
- 10.6k
- 派生
- 1.5k
- 平均合并
- 3 小时 57 分钟
- 30 天内合并 PR
- 4
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
abetlen/llama-cpp-python 的其他 Issue
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently drops可能已有人在做 @Belal0066 于 19 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
abetlen/llama-cpp-python#2371 ·
维护者通常 1 天内回复
-
uv add llama-cpp-python wheels fails for versions above 0.3.30可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2352 · 1 条评论 · 2 个 reaction ·
维护者通常 1 天内回复
-
Docs: consolidate build-from-source and GPU backend guide可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 75/100
abetlen/llama-cpp-python#2314 ·
维护者通常 1 天内回复
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_array可能重新可做 @lxcxjxhx 于 96 天前认领,目前没有进行中的 PR。 未关闭
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2211 · 2 条评论 ·
维护者通常 1 天内回复
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusingly可能已有人在做 @Anai-Guo 于 36 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2210 ·
维护者通常 1 天内回复
查看 abetlen/llama-cpp-python 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 83/100
PedestrianDynamics/pyFDS-Evac#766 ·
维护者通常 1 天内回复
-
难度 1/5 1-3 小时 新手友好度 91/100
alchaincyf/nuwa-skill#86 ·
-
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 2 天内回复
-
Docs Needs Triage
难度 1/5 1 小时以内 新手友好度 88/100
pandas-dev/pandas#71055 ·
维护者通常 1 天内回复
-
[Bug]: graphify reads files that git's global ignore file hides可能已有人在做 @smngvlkz 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 72/100
Graphify-Labs/graphify#4335 · 1 条评论 ·
维护者通常 1 天内回复