Can't compute multiple embeddings in a single call
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
- 難度
- 3/5
- 預估耗時
- 1-2 天
- 新手友好度
- 40/100
- Issue 類型
- 缺陷
- 描述清晰度
- 描述清楚
- 活躍度
- 停滯
- 領域
- ai, machine-learning
研究方向
錯誤發生在 llama_cpp/llama.py 約第 1108 行的 embed 方法以及第 1045 行的 decode_batch 中。先檢查 C++ 綁定(_internals.py)中的 batch 初始化和序列 ID 處理。查看 v0.3.14 與目前版本之間 embedding 邏輯的變更。執行提供的重現指令碼以查看確切的錯誤,然後檢查 llama.cpp 函式庫對多個序列進行 batch 解碼的實作。
由索引模型根據 Issue 內容生成。
描述
Prerequisites
Please answer the following questions for yourself before submitting an issue.
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
Running this code:
model = llama_cpp.Llama ("mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
embeddings = model.embed (["Hello", "World"])
used to work in v0.3.14
Current Behavior
The code raises an exception RuntimeError: llama_decode returned -1. The following messages are printed to the console:
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
Environment and Context
llama-cpp-python was compiled in CUDA mode
Failure Information (for bugs)
Please help provide information about the failure if this is a bug. If it is not a bug, please remove the rest of this template.
Steps to Reproduce
Python 3.11.2 (main, Apr 28 2025, 14:11:48) [GCC 12.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import llama_cpp
>>> model = llama_cpp.Llama ("../models/mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
...
>>> embeddings = model.embed (["Hello", "World"])
decode: cannot decode batches with this context (calling encode() instead)
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
llama_decode: failed to decode, ret = -1
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File ".../site-packages/llama_cpp/llama.py", line 1108, in embed
decode_batch(s_batch)
File ".../site-packages/llama_cpp/llama.py", line 1045, in decode_batch
self._ctx.decode(self._batch)
File ".../site-packages/llama_cpp/_internals.py", line 327, in decode
raise RuntimeError(f"llama_decode returned {return_code}")
RuntimeError: llama_decode returned -1
- 主要語言
- Python
- 星號
- 10.6k
- 分支
- 1.5k
- 平均合併
- 3 小時 57 分鐘
- 30 天內合併 PR
- 4
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
abetlen/llama-cpp-python 的其他 Issue
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently drops可能已有人在做 @Belal0066 於 15 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 88/100
abetlen/llama-cpp-python#2371 ·
維護者通常 1 天內回覆
-
uv add llama-cpp-python wheels fails for versions above 0.3.30可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2352 · 1 則留言 · 2 個 reaction ·
維護者通常 1 天內回覆
-
Docs: consolidate build-from-source and GPU backend guide可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 75/100
abetlen/llama-cpp-python#2314 ·
維護者通常 1 天內回覆
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_array可能重新可做 @lxcxjxhx 於 92 天前認領,目前沒有進行中的 PR。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2211 · 2 則留言 ·
維護者通常 1 天內回覆
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusingly可能已有人在做 @Anai-Guo 於 33 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 65/100
abetlen/llama-cpp-python#2210 ·
維護者通常 1 天內回覆
查看 abetlen/llama-cpp-python 的全部 Issue
相似的 Issue
-
enhancement
難度 2/5 1-3 小時 新手友好度 85/100
epam/ai-dial-quickapps-backend#628 ·
維護者通常 2 天內回覆
-
難度 1/5 1 小時以內 新手友好度 69/100
timqian/chinese-independent-blogs#2235 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 68/100
eclipse-score/coverage_tool#27 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 65/100
bojieli/ai-agent-book#1174 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 82/100
RedHatQE/mtv-api-tests#721 ·
維護者通常 1 天內回覆