Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_array
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 65/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- python
- Lĩnh vực
- backend-api-design
Hướng nghiên cứu
Issue nằm trong llama_cpp/llama.py, khoảng dòng 1678, nơi Llama.embed() gọi _batch.add_sequence với ba đối số. So sánh với lời gọi đúng gồm bốn đối số trong llama_cpp/llama_embedding.py, khoảng dòng 262. Cập nhật lời gọi để khớp với chữ ký, sử dụng mảng token, mảng vị trí, các ID chuỗi và mảng logits. Kiểm thử bằng cách chạy script tái hiện được cung cấp với một mô hình hỗ trợ embeddings.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Prerequisites
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
Llama.embed() should successfully compute embeddings when called on a model constructed with embeddings=True.
Current Behavior
Llama.embed() raises a TypeError immediately, before any embedding is computed:
TypeError: LlamaBatch.add_sequence() missing 1 required positional argument: 'logits_array'
The cause: Llama.embed() in llama_cpp/llama.py (around line 1678) calls add_sequence with three positional arguments:
self._batch.add_sequence(tokens, p_batch, logits_all)
But LlamaBatch.add_sequence in llama_cpp/_internals.py (around line 1013) requires four:
def add_sequence(
self,
token_array: Sequence[int],
pos_array: Sequence[int],
seq_ids: Sequence[Sequence[int]],
logits_array: Sequence[bool]
)
llama_cpp/llama_embedding.py (around line 262) already calls add_sequence correctly with the four-arg shape — the call site in Llama.embed() was apparently missed during the LlamaBatch.add_sequence refactor.
Environment and Context
- Hardware: x86_64, NVIDIA GeForce RTX 4090
- OS: Windows 10 22H2
- Python 3.12.9
- llama-cpp-python 0.3.36 (CUDA 12.8 prebuilt wheel)
$ python --version
Python 3.12.9
$ pip show llama-cpp-python | findstr Version
Version: 0.3.36
Failure Information (for bugs)
This is a clean regression — LlamaBatch.add_sequence was refactored from a 3-arg signature to a 4-arg one, and the call sites were updated everywhere except in Llama.embed(). llama_embedding.py shows what the new shape should look like for the embedding code path.
Steps to Reproduce
from llama_cpp import Llama
m = Llama(model_path="path/to/model.gguf", embeddings=True)
m.embed("hello")
Result:
TypeError: LlamaBatch.add_sequence() missing 1 required positional argument: 'logits_array'
Failure Logs
Traceback (most recent call last):
File "...\Lib\site-packages\llama_cpp\llama.py", line 1678, in embed
self._batch.add_sequence(tokens, p_batch, logits_all)
TypeError: LlamaBatch.add_sequence() missing 1 required positional argument: 'logits_array'
Suggested fix
Mirror the call shape already used in llama_cpp/llama_embedding.py:
# In llama.py Llama.embed(), replace:
self._batch.add_sequence(tokens, p_batch, logits_all)
# With something like:
self._batch.add_sequence(
token_array=tokens,
pos_array=list(range(len(tokens))),
seq_ids=[p_batch],
logits_array=[True] * len(tokens) if logits_all else [False] * (len(tokens) - 1) + [True],
)
Workaround
Monkey-patching LlamaBatch.add_sequence to detect 3-arg legacy calls and synthesize the missing pos_array works as a stopgap. Hit while running Tencent's HY-Motion text-to-motion model, whose text encoder uses Llama.embed() against GGUF Qwen3 weights.
- Ngôn ngữ chính
- Python
- Star
- 10.6k
- Fork
- 1.5k
- Merge trung bình
- 23 phút
- Pull request đã merge (30 ngày)
- 1
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của abetlen/llama-cpp-python
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2352 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2210 ·
-
Improve error messages Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2145 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 75/100
abetlen/llama-cpp-python#2135 · 4 reaction ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
abetlen/llama-cpp-python#2096 · 1 bình luận ·
Tất cả issue của abetlen/llama-cpp-python
Issue tương tự
-
triage/confirmed
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
agentscope-ai/agentscope#2775 ·
-
comp/desktop P3 type/bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
NousResearch/hermes-agent#118866 ·
-
bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
apache/cloudstack#14222 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100