long_term_memory: true without an embedding model crashes the entire server (unrecovered panic in saveCurrentConversation)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
Línea de trabajo
Comienza en knowledgebase.go, en Agent.saveCurrentConversation, y luego sigue el callback a través de consumeJob de agent.go y JobResult.Finish de result.go. Reproduce el problema con long_term_memory habilitado y sin un modelo de embeddings ni una base de conocimiento, y verifica que la solicitud mal configurada informe de un error sin terminar el servidor ni afectar a otros agentes.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Hi,
This is found and redacted by Claude, as I'm fucking unable to explain it better.
It's not a code review, but an actual crash I encountered while I let it work for me.
LocalAI version:
localai/localai:v4.9.0-gpu-nvidia-cuda-13 (LocalAI v4.9.0, commit f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e; vendored github.com/mudler/LocalAGI v0.0.0-20260606071251-14aed1ae4336)
Environment, CPU architecture, OS, and Version:
Linux ia 6.8.0-139-generic #139-Ubuntu SMP PREEMPT_DYNAMIC Sat Aug 1 03:52:05 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux
Bare-metal host (not a VM), Ubuntu 24.04.4 LTS, NVIDIA GeForce RTX 3090, driver 580.173.02, CUDA 13.0. Running via Docker Compose with runtime: nvidia. Backend: llama-cpp / cuda13-llama-cpp.
Describe the bug
Creating an agent via POST /api/agents (or PUT /api/agents/{name}) with "long_term_memory": true but without an embedding model installed/configured causes an unrecovered Go panic (nil pointer dereference) the first time that agent finishes a chat job and the framework tries to save the conversation. The panic is not scoped to the misconfigured agent's own goroutine — it crashes the entire local-ai process (PID 1 in the container). Docker's restart policy (unless-stopped) then restarts the container, which wipes the in-memory observables/job history for every agent on the instance (not just the misconfigured one) and briefly takes the whole API down for all users/agents while it reloads.
Root cause appears to be that long_term_memory depends on an embedding model (default seems to be granite-embedding-107m-multilingual, built-in chromem vector store) that was never installed on this instance, and the save-conversation code path dereferences it without a nil check.
To Reproduce
- Start LocalAI v4.9.0 with the agent pool enabled (default) and no embedding model installed (fresh instance, or one where
granite-embedding-107m-multilingualwas never pulled via the gallery). - Create an agent:
curl -X POST http://localhost:8080/api/agents \ -H "Content-Type: application/json" \ -d '{ "name": "test-agent", "model": "<any working chat model>", "system_prompt": "You are a test agent.", "api_url": "http://localhost:8080/v1", "api_key": "sk-local", "long_term_memory": true }' - Send it a chat message:
curl -X POST http://localhost:8080/api/agents/test-agent/chat \ -H "Content-Type: application/json" \ -d '{"message": "hello"}' - Wait for the model to generate a response and for the framework to attempt saving the conversation — the whole server crashes and restarts a few seconds later.
Expected behavior
The API should return an error for that specific request/agent (e.g. "long_term_memory requires an embedding model / knowledge base to be configured"), and the rest of the server (other agents, in-flight requests) should be unaffected. At minimum, the panic should be recovered at the job/goroutine level so one misconfigured agent can't take down the whole process.
Logs
Saving conversation agent="supervisor" conversation size=4
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x30 pc=0xf9a3a9]
goroutine 7150 [running]:
github.com/mudler/LocalAGI/core/agent.(*Agent).saveCurrentConversation(0x1c1dad1543c0, {0x1c1da66b8f08, 0x4, 0x1c1da591c730?})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/knowledgebase.go:176 +0x629
github.com/mudler/LocalAGI/core/agent.(*Agent).consumeJob.func8({0x1c1da66b8f08?, 0x4?, 0x4?})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1420 +0x28
github.com/mudler/LocalAGI/core/types.(*JobResult).Finish(0x1c1da50fd290, {0x0?, 0x0?})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/types/result.go:43 +0xdf
github.com/mudler/LocalAGI/core/agent.(*Agent).consumeJob(0x1c1dad1543c0, 0x1c1da566e180, {0x48076fd, 0x4})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1423 +0x2b27
github.com/mudler/LocalAGI/core/agent.(*Agent).run(0x1c1dad1543c0, 0x1c1da4854e70)
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1543 +0xfd
github.com/mudler/LocalAGI/core/agent.(*Agent).Run.func1()
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1521 +0x3a
created by github.com/mudler/LocalAGI/core/agent.(*Agent).Run in goroutine 6913
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1520 +0x1db
docker inspect on the container right after the crash showed ExitCode=0, OOMKilled=false, Error="", and docker events showed a die event with no preceding kill/stop action — confirming this is the process self-terminating on the panic, not an external kill/OOM.
Additional context
Workaround that resolves it: install an embedding model first (POST /models/apply with {"id": "localai@granite-embedding-107m-multilingual"}), then set both long_term_memory: true and enable_kb: true together on the agent (long_term_memory alone is not enough / is what triggers the crash). Retested with the same repro steps after applying the workaround: no panic, container RestartCount unchanged, GET /api/agents/collections shows a collection for the agent, and the job's observables history shows a clean "Recall" step (KB auto-search) before the actual response.
Found while building a multi-agent setup (Planner/Coder/Reviewer/Supervisor) on top of LocalAI's native agent pool.
Regards
- Lenguaje dominante
- Go
- Estrellas
- 49.2k
- Forks
- 4.5k
- Merge medio
- 1 d 7 h
- PR fusionados (30 d)
- 357
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de mudler/LocalAI
-
bug unconfirmed
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
mudler/LocalAI#12337 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
mudler/LocalAI#11995 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
mudler/LocalAI#11991 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
mudler/LocalAI#11348 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
bug unconfirmed
Dificultad 4/5 3-5 días Aptitud para principiantes 70/100
mudler/LocalAI#12331 · 1 comentario ·
Los mantenedores suelen responder en 1 día
Todos los issues de mudler/LocalAI
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
gruntwork-io/boilerplate#329 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
prime-radiant-inc/evener#3291 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Netcracker/qubership-apihub-backend#582 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 1 día