Idle watchdog reaps a stale address and shuts down a different, live backend
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 70/100
Direzione di ricerca
Start with pkg/model/watchdog.go and pkg/model/process.go, tracing WatchDog.Add, addressModelMap, checkIdle/checkBusy, and deleteProcess. Verify the fix by ensuring externally stopped addresses leave no stale watchdog state and that an expired old address cannot shut down a replacement backend; cover the behavior with regression tests around the watchdog lifecycle.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Written on behalf of YourYoungerBrothersPug by Claude Code.
LocalAI version:
Observed on v4.9.0 (localai/localai:v4.9.0-gpu-nvidia-cuda-12). The code path is identical in v4.10.0 and on
master @ 590512d9: pkg/model/watchdog.go has not changed between them.
Environment, CPU architecture, OS, and Version:
Linux 6.6.114.1-microsoft-standard-WSL2 #1 SMP PREEMPT_DYNAMIC Mon Dec 1 20:46:23 UTC 2025 x86_64 GNU/Linux
(uname -ainside the LocalAI container)- x86_64, Docker on WSL2, NVIDIA RTX 3060 (12 GB), CUDA 12
- Single node. Not distributed mode: no workers, no federation.
Describe the bug
WatchDog is told when a backend starts but never when one stops, unless the watchdog itself stopped it.
Any backend torn down by another path (POST /backend/shutdown, ShutdownModel, a crash) leaves its address
behind in idleTime and addressModelMap permanently.
One idle timeout later checkIdle() fires on that dead address, resolves it to a model name, and calls
ShutdownModel(model). That stops whatever process is serving that model at that moment, which is a
different, healthy backend, and often one that is still loading.
Registration happens on load, in pkg/model/process.go:
ml.wd.Add(serverAddress, grpcControlProcess)
ml.wd.AddAddressModelMap(serverAddress, id)
There is no deregistration. untrack(address) is unexported and is called from four places, all in watchdog.go
and all watchdog-initiated evictions: checkIdle, checkBusy, collectEvictionsLocked, evictLRUModel.
deleteProcess (pkg/model/process.go) runs the unload hooks, calls Free(), stops the process, does
store.Delete(s) and cleanupProcessRuntime(process), and touches no watchdog state at all. Grepping the whole
pkg/model package for wd. finds only Add, AddAddressModelMap, RegisterModelSize, UpdateLastUsed,
EnforceLRULimit and EnforceGroupExclusivity: nothing that removes an address.
An address enters idleTime in finishRequestLocked (wd.idleTime[address] = now), so any backend that served
at least one request and was then stopped externally has a stale entry that will eventually come due. Then, in
checkIdle():
for address, t := range wd.idleTime {
if time.Since(t) > wd.idletimeout {
model, ok := wd.addressModelMap[address]
...
modelsToShutdown = append(modelsToShutdown, model)
wd.untrack(address)
}
}
...
for _, model := range modelsToShutdown {
if err := wd.pm.ShutdownModel(model); err != nil { ... }
}
The timer belongs to the address; the shutdown names the model. Nothing checks that the address being
reaped is still the address currently serving that model.
To Reproduce
- Run a single node with
LOCALAI_WATCHDOG_IDLE=true. SetLOCALAI_WATCHDOG_IDLE_TIMEOUT=1mto see it within a
minute; the 15m default behaves identically, just later. - Send one request to model
M, so a backend starts and anidleTimeentry is created for its address. - Stop it externally:
POST /backend/shutdown {"model": "M"}. - Send another request to
M. A new backend starts on a new address. Confirm the two differ in the
BackendLoader startinglines. - Wait out the idle timeout measured from step 2, keeping
Min use so the new backend never goes idle itself.
The watchdog reaps the dead address from step 2 and calls ShutdownModel("M"), stopping the healthy backend
from step 4. Any workload that stops backends deliberately arms one of these per stop, so they accumulate.
Expected behavior
An address that no longer has a running backend should not be reaped, and reaping one should never stop a
different, live process. Concretely: after step 3, the stale entry should be gone, and step 5 should be a no-op.
Logs
Observed ten times in three hours on one box, each exactly the 15m timeout after a different, already-dead
address for that model had stopped. The worst case is a kill landing during a load, because a load is the longest
thing that happens to a model. Trimmed to the relevant lines (these appear at info level; the first line is an
annotation, not log output):
[a request for granite-4.1-8b arrives]
22:08:42 BackendLoader starting modelID=granite-4.1-8b -> 127.0.0.1:38831
22:09:11 [WatchDog] Address is idle for too long, killing it address=127.0.0.1:45301
22:09:32 Backend process stopped id=granite-4.1-8b address=127.0.0.1:38831 exitCode=-1
22:09:32 POST /v1/chat/completions 500 dial tcp 127.0.0.1:38831: connect: connection refused
127.0.0.1:45301 had been dead since 21:54:12, exactly 15m before the reap. 127.0.0.1:38831 was 30 seconds
into a healthy load and is what actually died. The client gets a 500 with nothing in the log connecting it to the
address that was reaped.
Additional context
Two things make this hard to spot from the logs:
- The
"Address is idle for too long, killing it"warning logs only the address. It never names the model it
is about to stop, nor the live address that will actually die. "[watchdog] error shutting down model: model not found"is the same bug, firing on a stale entry whose model
happens not to be loaded at that moment. It reads like a harmless line and is the clearest early tell that stale
entries are accumulating.
There is also a secondary effect: untrack does delete(wd.modelSizes, modelID) for the model resolved from the
stale address, so reaping a dead address also discards the live model's registered size, quietly degrading
size-aware eviction.
Notes:
- A longer idle timeout is not a fix. The duration is not the fault; any timeout arms the same kill, only later.
- Disabling the idle watchdog avoids it, at the cost of the keep-alive behaviour it exists to provide.
checkBusyreads the sameaddressModelMapand is worth checking for the same stale-entry exposure.
Suggested fix: untrack the address when a backend stops for any reason, not only when the watchdog stopped it,
e.g. have deleteProcess notify the watchdog (an exported Untrack(address), or a ModelUnloadHook), using the
address recorded for that model at load time. Belt and braces: in checkIdle, before shutting down, confirm the
address being reaped is still the one associated with the currently-loaded process for that model, and skip the
shutdown if it is not.
- Lingua principale
- Go
- Stelle
- 49.2k
- Fork
- 4.5k
- Merge medio
- 1g 7h
- PR unite (30g)
- 357
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di mudler/LocalAI
-
bug unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
mudler/LocalAI#12337 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11995 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
mudler/LocalAI#11991 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasApertaenhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11348 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
feat: add automatic MCP transport selection for 2024-11-05 / 2025-03-26 / 2025-06-18 vs 2025-11-25Apertaenhancement
Difficoltà 3/5 1-2 giorni Idoneità per principianti 65/100
mudler/LocalAI#12262 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di mudler/LocalAI
Issue simili
-
agentic-workflows
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
priority/4/normal status/needs-triage type/bug/unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
authelia/authelia#13292 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
blinklabs-io/actions#138 ·
I maintainer di solito rispondono entro 1 giorno
-
[UI] AlbumDetails collapses multi-genre list to single primary genre on viewports < lg breakpointAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno