Multiple calls to create_chat_completion() fail with "llama_decode: failed to decode, ret = -1"
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 3/5
- Temps estimé
- 1-2 jours
- Accessibilité débutants
- 55/100
- Type d'issue
- Bug
- Clarté
- Clairement spécifiée
- Activité
- À l'abandon
- Stack technique
- python
- Domaine
- ai, backend-api-design
Piste de recherche
L’issue se trouve dans la méthode create_chat_completion de la classe llama_cpp.Llama. Examinez le code source de cette méthode et les bindings C++ sous-jacents pour voir comment le contexte et le cache KV sont gérés entre les appels. L’utilisateur a constaté que l’appel de llm.reset() ou de llm._ctx.kv_cache_clear() fonctionne ; la correction implique donc probablement de réinitialiser automatiquement l’état après une completion. Vérifiez la suite de tests pour trouver des schémas similaires. « Terminé » signifie que deux appels consécutifs à create_chat_completion réussissent sans réinitialisation manuelle.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Prerequisites
Please answer the following questions for yourself before submitting an issue.
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
I am running LiquidAI's LFM2.5-1.2B-Instruct model. Calling create_chat_completion() multiple times should not throw error.
Current Behavior
When trying to call create_chat_completion twice in a row, the model throws error "llama_decode: failed to decode, ret = -1"
Environment and Context
I am using llama-cpp-python v0.3.16. Python version is 3.13.9.
- Windows 11
Failure Information (for bugs)
On tracing back the issue, it looks like there needs to be a context and cache reset after each chat_completion call, which isn't happening yet.
Steps to Reproduce
from pathlib import Path
from llama_cpp import Llama
llm = Llama(
model_path=str(Path.home() / "AppData/Local/llama.cpp/LiquidAI_LFM2.5-1.2B-Instruct-GGUF_LFM2.5-1.2B-Instruct-Q4_K_M.gguf"),
n_ctx=1000
)
system_prompt = """
\nYou are a helpful assistant
"""
prompt = """
suggest me places to visit during winter season
"""
response = llm.create_chat_completion(
messages = [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)
# llm.reset() # Using this works
# llm._ctx.kv_cache_clear() # Using this works
response = llm.create_chat_completion(
messages = [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)
Failure Logs
init: the tokens of sequence 0 in the input batch have inconsistent sequence positions:
- the last position stored in the memory module of the context (i.e. the KV cache) for sequence 0 is X = 519
- the tokens for sequence 0 in the input batch have a starting position of Y = 29
it is required that the sequence positions remain consecutive: Y = X + 1
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
- Langage dominant
- Python
- Étoiles
- 10.6k
- Forks
- 1.5k
- Merge moyen
- 3 h 57 min
- PR mergées (30 j)
- 4
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Aucun modèle de pull request
- Lire le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de abetlen/llama-cpp-python
-
Seven llama_sampler_init_* bindings admit keyword arguments that the ctypes function object silently dropsPeut-être pris @Belal0066 l’a pris il y a 19 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 88/100
abetlen/llama-cpp-python#2371 ·
Les mainteneurs répondent en général sous 1 jour
-
uv add llama-cpp-python wheels fails for versions above 0.3.30Peut-être pris Une pull request liée à cette issue est ouverte ou déjà fusionnée. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
abetlen/llama-cpp-python#2352 · 1 commentaire · 2 réactions ·
Les mainteneurs répondent en général sous 1 jour
-
Docs: consolidate build-from-source and GPU backend guidePeut-être pris Une pull request liée à cette issue est ouverte ou déjà fusionnée. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
abetlen/llama-cpp-python#2314 ·
Les mainteneurs répondent en général sous 1 jour
-
Llama.embed() calls LlamaBatch.add_sequence with old 3-arg signature; missing logits_arrayPeut-être à nouveau libre @lxcxjxhx l’a pris il y a 95 jours, et aucune pull request n’est ouverte. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
abetlen/llama-cpp-python#2211 · 2 commentaires ·
Les mainteneurs répondent en général sous 1 jour
-
Llama() silently accepts and discards `embedding` kwarg; .embed() then raises confusinglyPeut-être pris @Anai-Guo l’a pris il y a 36 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
abetlen/llama-cpp-python#2210 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de abetlen/llama-cpp-python
Issues similaires
-
feedback simulation workshop
Difficulté 2/5 1-3 heures Accessibilité débutants 73/100
githubnext/gh-aw-workshop#4455 ·
Les mainteneurs répondent en général sous 1 jour
-
Triage 🩺
Difficulté 2/5 1-3 heures Accessibilité débutants 76/100
Les mainteneurs répondent en général sous 1 jour
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOuverteneeds-triage
Difficulté 2/5 1-3 heures Accessibilité débutants 77/100
krkn-chaos/krkn#1627 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
NousResearch/hermes-agent#136483 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 1/5 Moins d'une heure Accessibilité débutants 88/100
Les mainteneurs répondent en général sous 1 jour