Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Multiple calls to create_chat_completion() fail with "llama_decode: failed to decode, ret = -1"

未关闭
#2,140 0 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
55/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
停滞
技术栈
python

调研方向

该 issue 位于 llama_cpp.Llama 类的 create_chat_completion 方法中。查看该方法的源代码以及底层的 C++ bindings,了解 context 和 KV cache 在多次调用之间是如何管理的。用户发现调用 llm.reset() 或 llm._ctx.kv_cache_clear() 有效,因此修复可能涉及在一次 completion 后自动重置 state。检查 test suite 中是否有类似模式。“完成”的标准是:连续两次调用 create_chat_completion 都能成功,且无需手动 reset。

由索引模型根据 Issue 内容生成。

描述

Prerequisites

Please answer the following questions for yourself before submitting an issue.

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Expected Behavior

I am running LiquidAI's LFM2.5-1.2B-Instruct model. Calling create_chat_completion() multiple times should not throw error.

Current Behavior

When trying to call create_chat_completion twice in a row, the model throws error "llama_decode: failed to decode, ret = -1"

Environment and Context

I am using llama-cpp-python v0.3.16. Python version is 3.13.9.

  • Windows 11

Failure Information (for bugs)

On tracing back the issue, it looks like there needs to be a context and cache reset after each chat_completion call, which isn't happening yet.

Steps to Reproduce


from pathlib import Path
from llama_cpp import Llama
llm = Llama(
    model_path=str(Path.home() / "AppData/Local/llama.cpp/LiquidAI_LFM2.5-1.2B-Instruct-GGUF_LFM2.5-1.2B-Instruct-Q4_K_M.gguf"),
    n_ctx=1000
)

system_prompt = """
\nYou are a helpful assistant
"""
prompt = """
suggest me places to visit during winter season
"""
response = llm.create_chat_completion(
      messages =  [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)
# llm.reset()                                               # Using this works
# llm._ctx.kv_cache_clear()                        # Using this works
response = llm.create_chat_completion(
      messages =  [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)

Failure Logs

init: the tokens of sequence 0 in the input batch have inconsistent sequence positions:
 - the last position stored in the memory module of the context (i.e. the KV cache) for sequence 0 is X = 519
 - the tokens for sequence 0 in the input batch have a starting position of Y = 29
 it is required that the sequence positions remain consecutive: Y = X + 1
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
主要语言
Python
星标
10.6k
派生
1.5k
平均合并
3 小时 57 分钟
30 天内合并 PR
4

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

abetlen/llama-cpp-python 的其他 Issue

查看 abetlen/llama-cpp-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。