Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Multiple calls to create_chat_completion() fail with "llama_decode: failed to decode, ret = -1"

Abierto
#2,140 0 comentarios 1 reacción 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
3/5
Tiempo estimado
1-2 días
Aptitud para principiantes
55/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Estancado
Stack tecnológico
python

Línea de trabajo

El issue está en el método create_chat_completion de la clase llama_cpp.Llama. Revisa el código fuente de ese método y los bindings subyacentes de C++ para ver cómo se gestionan el contexto y la caché KV entre llamadas. El usuario descubrió que llamar a llm.reset() o llm._ctx.kv_cache_clear() funciona, por lo que la solución probablemente implique restablecer automáticamente el estado después de una completion. Revisa la suite de tests en busca de patrones similares. «Hecho» es cuando dos llamadas consecutivas a create_chat_completion tienen éxito sin un restablecimiento manual.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Prerequisites

Please answer the following questions for yourself before submitting an issue.

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Expected Behavior

I am running LiquidAI's LFM2.5-1.2B-Instruct model. Calling create_chat_completion() multiple times should not throw error.

Current Behavior

When trying to call create_chat_completion twice in a row, the model throws error "llama_decode: failed to decode, ret = -1"

Environment and Context

I am using llama-cpp-python v0.3.16. Python version is 3.13.9.

  • Windows 11

Failure Information (for bugs)

On tracing back the issue, it looks like there needs to be a context and cache reset after each chat_completion call, which isn't happening yet.

Steps to Reproduce


from pathlib import Path
from llama_cpp import Llama
llm = Llama(
    model_path=str(Path.home() / "AppData/Local/llama.cpp/LiquidAI_LFM2.5-1.2B-Instruct-GGUF_LFM2.5-1.2B-Instruct-Q4_K_M.gguf"),
    n_ctx=1000
)

system_prompt = """
\nYou are a helpful assistant
"""
prompt = """
suggest me places to visit during winter season
"""
response = llm.create_chat_completion(
      messages =  [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)
# llm.reset()                                               # Using this works
# llm._ctx.kv_cache_clear()                        # Using this works
response = llm.create_chat_completion(
      messages =  [{"role": "system", "content": system_prompt}, {"role": "user", "content": prompt}],
)
print(response)

Failure Logs

init: the tokens of sequence 0 in the input batch have inconsistent sequence positions:
 - the last position stored in the memory module of the context (i.e. the KV cache) for sequence 0 is X = 519
 - the tokens for sequence 0 in the input batch have a starting position of Y = 29
 it is required that the sequence positions remain consecutive: Y = X + 1
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
Lenguaje dominante
Python
Estrellas
10.6k
Forks
1.5k
Merge medio
3 h 57 min
PR fusionados (30 d)
4

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de abetlen/llama-cpp-python

Todos los issues de abetlen/llama-cpp-python

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.