long_term_memory: true without an embedding model crashes the entire server (unrecovered panic in saveCurrentConversation)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 68/100
Direzione di ricerca
Inizia in knowledgebase.go, in Agent.saveCurrentConversation, quindi segui il callback attraverso consumeJob in agent.go e JobResult.Finish in result.go. Riproduci il problema con long_term_memory abilitato e senza un modello di embedding né una knowledge base, quindi verifica che la richiesta configurata in modo errato segnali un errore senza terminare il server né influire sugli altri agenti.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi,
This is found and redacted by Claude, as I'm fucking unable to explain it better.
It's not a code review, but an actual crash I encountered while I let it work for me.
LocalAI version:
localai/localai:v4.9.0-gpu-nvidia-cuda-13 (LocalAI v4.9.0, commit f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e; vendored github.com/mudler/LocalAGI v0.0.0-20260606071251-14aed1ae4336)
Environment, CPU architecture, OS, and Version:
Linux ia 6.8.0-139-generic #139-Ubuntu SMP PREEMPT_DYNAMIC Sat Aug 1 03:52:05 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux
Bare-metal host (not a VM), Ubuntu 24.04.4 LTS, NVIDIA GeForce RTX 3090, driver 580.173.02, CUDA 13.0. Running via Docker Compose with runtime: nvidia. Backend: llama-cpp / cuda13-llama-cpp.
Describe the bug
Creating an agent via POST /api/agents (or PUT /api/agents/{name}) with "long_term_memory": true but without an embedding model installed/configured causes an unrecovered Go panic (nil pointer dereference) the first time that agent finishes a chat job and the framework tries to save the conversation. The panic is not scoped to the misconfigured agent's own goroutine — it crashes the entire local-ai process (PID 1 in the container). Docker's restart policy (unless-stopped) then restarts the container, which wipes the in-memory observables/job history for every agent on the instance (not just the misconfigured one) and briefly takes the whole API down for all users/agents while it reloads.
Root cause appears to be that long_term_memory depends on an embedding model (default seems to be granite-embedding-107m-multilingual, built-in chromem vector store) that was never installed on this instance, and the save-conversation code path dereferences it without a nil check.
To Reproduce
- Start LocalAI v4.9.0 with the agent pool enabled (default) and no embedding model installed (fresh instance, or one where
granite-embedding-107m-multilingualwas never pulled via the gallery). - Create an agent:
curl -X POST http://localhost:8080/api/agents \ -H "Content-Type: application/json" \ -d '{ "name": "test-agent", "model": "<any working chat model>", "system_prompt": "You are a test agent.", "api_url": "http://localhost:8080/v1", "api_key": "sk-local", "long_term_memory": true }' - Send it a chat message:
curl -X POST http://localhost:8080/api/agents/test-agent/chat \ -H "Content-Type: application/json" \ -d '{"message": "hello"}' - Wait for the model to generate a response and for the framework to attempt saving the conversation — the whole server crashes and restarts a few seconds later.
Expected behavior
The API should return an error for that specific request/agent (e.g. "long_term_memory requires an embedding model / knowledge base to be configured"), and the rest of the server (other agents, in-flight requests) should be unaffected. At minimum, the panic should be recovered at the job/goroutine level so one misconfigured agent can't take down the whole process.
Logs
Saving conversation agent="supervisor" conversation size=4
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x30 pc=0xf9a3a9]
goroutine 7150 [running]:
github.com/mudler/LocalAGI/core/agent.(*Agent).saveCurrentConversation(0x1c1dad1543c0, {0x1c1da66b8f08, 0x4, 0x1c1da591c730?})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/knowledgebase.go:176 +0x629
github.com/mudler/LocalAGI/core/agent.(*Agent).consumeJob.func8({0x1c1da66b8f08?, 0x4?, 0x4?})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1420 +0x28
github.com/mudler/LocalAGI/core/types.(*JobResult).Finish(0x1c1da50fd290, {0x0?, 0x0?})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/types/result.go:43 +0xdf
github.com/mudler/LocalAGI/core/agent.(*Agent).consumeJob(0x1c1dad1543c0, 0x1c1da566e180, {0x48076fd, 0x4})
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1423 +0x2b27
github.com/mudler/LocalAGI/core/agent.(*Agent).run(0x1c1dad1543c0, 0x1c1da4854e70)
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1543 +0xfd
github.com/mudler/LocalAGI/core/agent.(*Agent).Run.func1()
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1521 +0x3a
created by github.com/mudler/LocalAGI/core/agent.(*Agent).Run in goroutine 6913
/root/go/pkg/mod/github.com/mudler/[email protected]/core/agent/agent.go:1520 +0x1db
docker inspect on the container right after the crash showed ExitCode=0, OOMKilled=false, Error="", and docker events showed a die event with no preceding kill/stop action — confirming this is the process self-terminating on the panic, not an external kill/OOM.
Additional context
Workaround that resolves it: install an embedding model first (POST /models/apply with {"id": "localai@granite-embedding-107m-multilingual"}), then set both long_term_memory: true and enable_kb: true together on the agent (long_term_memory alone is not enough / is what triggers the crash). Retested with the same repro steps after applying the workaround: no panic, container RestartCount unchanged, GET /api/agents/collections shows a collection for the agent, and the job's observables history shows a clean "Recall" step (KB auto-search) before the actual response.
Found while building a multi-agent setup (Planner/Coder/Reviewer/Supervisor) on top of LocalAI's native agent pool.
Regards
- Lingua principale
- Go
- Stelle
- 49.2k
- Fork
- 4.5k
- Merge medio
- 1g 7h
- PR unite (30g)
- 362
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di mudler/LocalAI
-
bug unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
mudler/LocalAI#12337 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11995 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
mudler/LocalAI#11991 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasApertaenhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11348 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
barholeurApertabug unconfirmed
Difficoltà 5/5 Più di una settimana Idoneità per principianti 10/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di mudler/LocalAI
Issue simili
-
[Docs] - Document minimum Terraform/OpenTofu version (>= 1.11) required by write-only argumentsAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 92/100
MagaluCloud/terraform-provider-mgc#323 ·
I maintainer di solito rispondono entro 11 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
rossoctl/context-guru#366 ·
I maintainer di solito rispondono entro 1 giorno
-
stage-fail
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
siyuan-note/bazaar#2293 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
piraeusdatastore/piraeus-operator#1070 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno