Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Can't compute multiple embeddings in a single call

未關閉
#2,051 4 則留言 5 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
3/5
預估耗時
1-2 天
新手友好度
40/100
Issue 類型
缺陷
描述清晰度
描述清楚
活躍度
停滯
技術堆疊
c, python

研究方向

錯誤發生在 llama_cpp/llama.py 約第 1108 行的 embed 方法以及第 1045 行的 decode_batch 中。先檢查 C++ 綁定(_internals.py)中的 batch 初始化和序列 ID 處理。查看 v0.3.14 與目前版本之間 embedding 邏輯的變更。執行提供的重現指令碼以查看確切的錯誤,然後檢查 llama.cpp 函式庫對多個序列進行 batch 解碼的實作。

由索引模型根據 Issue 內容生成。

描述

Prerequisites

Please answer the following questions for yourself before submitting an issue.

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Expected Behavior

Running this code:

model = llama_cpp.Llama ("mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
embeddings = model.embed (["Hello", "World"])

used to work in v0.3.14

Current Behavior

The code raises an exception RuntimeError: llama_decode returned -1. The following messages are printed to the console:

init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch

Environment and Context

llama-cpp-python was compiled in CUDA mode

Failure Information (for bugs)

Please help provide information about the failure if this is a bug. If it is not a bug, please remove the rest of this template.

Steps to Reproduce

Python 3.11.2 (main, Apr 28 2025, 14:11:48) [GCC 12.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import llama_cpp
>>> model = llama_cpp.Llama ("../models/mxbai-embed-xsmall-v1-q8_0.gguf", embedding = True)
...
>>> embeddings = model.embed (["Hello", "World"])
decode: cannot decode batches with this context (calling encode() instead)
init: invalid seq_id[3][0] = 1 >= 1
encode: failed to initialize batch
llama_decode: failed to decode, ret = -1
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File ".../site-packages/llama_cpp/llama.py", line 1108, in embed
    decode_batch(s_batch)
  File ".../site-packages/llama_cpp/llama.py", line 1045, in decode_batch
    self._ctx.decode(self._batch)
  File ".../site-packages/llama_cpp/_internals.py", line 327, in decode
    raise RuntimeError(f"llama_decode returned {return_code}")
RuntimeError: llama_decode returned -1
主要語言
Python
星號
10.6k
分支
1.5k
平均合併
3 小時 57 分鐘
30 天內合併 PR
4

環境準備

  • 沒有 Dockerfile 或 Docker Compose 檔案
  • 沒有 Pull Request 範本
  • 閱讀貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

abetlen/llama-cpp-python 的其他 Issue

查看 abetlen/llama-cpp-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。