Regression in unified KV cache appears after `llama.cpp` release b5912 in b5913
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 35/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
调研方向
该 issue 指出了 llama.cpp 的统一 KV cache(src/llama-kv-cache-unified.cpp 第 222 行)中的一个回归问题,该问题影响 Python bindings。首先使用 llama-cpp-python 基于 llama.cpp 的提交 b5912 和 b5913 重现该故障。检查这两个提交之间 llama.cpp 在 KV cache 处理方面的变更,重点关注 sequence ID 管理和 buffer 分离。修复可能需要理解 C++/Python 接口以及 KV cache 的内部状态。
由索引模型根据 Issue 内容生成。
描述
This issue concerns the llama-cpp-python community but was filed on the llama.cpp tracker first: https://github.com/ggml-org/llama.cpp/issues/14847.
I just wanted to bring it to your attention. I can relocate the issue if it is more relevant here. For your convenience, the issue description is reproduced here:
Running llama-cpp-python against llama.cpp compiled after b5912, in b5913, results in:
llama.cpp/src/llama-kv-cache-unified.cpp:222: GGML_ASSERT(seq_id >= 0 && (size_t) seq_id < seq_to_stream.size()) failed
It appears to be a regression in sequence ID handling or unified KV cache logic affecting external bindings. This is consistent with the heavy work done on the kv-cache to prepare K/V buffers for separation in b5913.
NOTE: llama-cli runs successfully, but running llama-cpp-python against llama.cpp with the same model fails.
- 主要语言
- Python
- 星标
- 10.6k
- 派生
- 1.5k
- 平均合并
- 3 小时 57 分钟
- 30 天内合并 PR
- 4
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
abetlen/llama-cpp-python 的其他 Issue
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently drops可能已有人在做 @Belal0066 于 15 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
abetlen/llama-cpp-python#2371 ·
维护者通常 1 天内回复
-
uv add llama-cpp-python wheels fails for versions above 0.3.30可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2352 · 1 条评论 · 2 个 reaction ·
维护者通常 1 天内回复
-
Docs: consolidate build-from-source and GPU backend guide可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 75/100
abetlen/llama-cpp-python#2314 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2211 · 2 条评论 ·
维护者通常 1 天内回复
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusingly可能已有人在做 @Anai-Guo 于 32 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2210 ·
维护者通常 1 天内回复
查看 abetlen/llama-cpp-python 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 92/100
-
Harmony OPeNDAP SubSetter (HOSS) Geographic LARC_CLOUD PREFIRE_SAT2_AUX-SAT R01 production
难度 2/5 1-3 小时 新手友好度 68/100
nasa/harmony-autotester#245 ·
-
enhancement
难度 2/5 1-3 小时 新手友好度 68/100
Deltares/imod-python#1928 ·
-
难度 1/5 1 小时以内 新手友好度 88/100
维护者通常 1 天内回复
-
feature
难度 2/5 1-3 小时 新手友好度 66/100