Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Idle watchdog reaps a stale address and shuts down a different, live backend

Aperta
#12,331 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
70/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
go
Ambito
backend

Direzione di ricerca

Start with pkg/model/watchdog.go and pkg/model/process.go, tracing WatchDog.Add, addressModelMap, checkIdle/checkBusy, and deleteProcess. Verify the fix by ensuring externally stopped addresses leave no stale watchdog state and that an expired old address cannot shut down a replacement backend; cover the behavior with regression tests around the watchdog lifecycle.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug unconfirmed

Written on behalf of YourYoungerBrothersPug by Claude Code.

LocalAI version:

Observed on v4.9.0 (localai/localai:v4.9.0-gpu-nvidia-cuda-12). The code path is identical in v4.10.0 and on
master @ 590512d9: pkg/model/watchdog.go has not changed between them.

Environment, CPU architecture, OS, and Version:

  • Linux 6.6.114.1-microsoft-standard-WSL2 #1 SMP PREEMPT_DYNAMIC Mon Dec 1 20:46:23 UTC 2025 x86_64 GNU/Linux
    (uname -a inside the LocalAI container)
  • x86_64, Docker on WSL2, NVIDIA RTX 3060 (12 GB), CUDA 12
  • Single node. Not distributed mode: no workers, no federation.

Describe the bug

WatchDog is told when a backend starts but never when one stops, unless the watchdog itself stopped it.
Any backend torn down by another path (POST /backend/shutdown, ShutdownModel, a crash) leaves its address
behind in idleTime and addressModelMap permanently.

One idle timeout later checkIdle() fires on that dead address, resolves it to a model name, and calls
ShutdownModel(model). That stops whatever process is serving that model at that moment, which is a
different, healthy backend, and often one that is still loading.

Registration happens on load, in pkg/model/process.go:

ml.wd.Add(serverAddress, grpcControlProcess)
ml.wd.AddAddressModelMap(serverAddress, id)

There is no deregistration. untrack(address) is unexported and is called from four places, all in watchdog.go
and all watchdog-initiated evictions: checkIdle, checkBusy, collectEvictionsLocked, evictLRUModel.
deleteProcess (pkg/model/process.go) runs the unload hooks, calls Free(), stops the process, does
store.Delete(s) and cleanupProcessRuntime(process), and touches no watchdog state at all. Grepping the whole
pkg/model package for wd. finds only Add, AddAddressModelMap, RegisterModelSize, UpdateLastUsed,
EnforceLRULimit and EnforceGroupExclusivity: nothing that removes an address.

An address enters idleTime in finishRequestLocked (wd.idleTime[address] = now), so any backend that served
at least one request and was then stopped externally has a stale entry that will eventually come due. Then, in
checkIdle():

for address, t := range wd.idleTime {
    if time.Since(t) > wd.idletimeout {
        model, ok := wd.addressModelMap[address]
        ...
        modelsToShutdown = append(modelsToShutdown, model)
        wd.untrack(address)
    }
}
...
for _, model := range modelsToShutdown {
    if err := wd.pm.ShutdownModel(model); err != nil { ... }
}

The timer belongs to the address; the shutdown names the model. Nothing checks that the address being
reaped is still the address currently serving that model.

To Reproduce

  1. Run a single node with LOCALAI_WATCHDOG_IDLE=true. Set LOCALAI_WATCHDOG_IDLE_TIMEOUT=1m to see it within a
    minute; the 15m default behaves identically, just later.
  2. Send one request to model M, so a backend starts and an idleTime entry is created for its address.
  3. Stop it externally: POST /backend/shutdown {"model": "M"}.
  4. Send another request to M. A new backend starts on a new address. Confirm the two differ in the
    BackendLoader starting lines.
  5. Wait out the idle timeout measured from step 2, keeping M in use so the new backend never goes idle itself.

The watchdog reaps the dead address from step 2 and calls ShutdownModel("M"), stopping the healthy backend
from step 4. Any workload that stops backends deliberately arms one of these per stop, so they accumulate.

Expected behavior

An address that no longer has a running backend should not be reaped, and reaping one should never stop a
different, live process. Concretely: after step 3, the stale entry should be gone, and step 5 should be a no-op.

Logs

Observed ten times in three hours on one box, each exactly the 15m timeout after a different, already-dead
address for that model had stopped. The worst case is a kill landing during a load, because a load is the longest
thing that happens to a model. Trimmed to the relevant lines (these appear at info level; the first line is an
annotation, not log output):

          [a request for granite-4.1-8b arrives]
22:08:42  BackendLoader starting  modelID=granite-4.1-8b  ->  127.0.0.1:38831
22:09:11  [WatchDog] Address is idle for too long, killing it  address=127.0.0.1:45301
22:09:32  Backend process stopped  id=granite-4.1-8b  address=127.0.0.1:38831  exitCode=-1
22:09:32  POST /v1/chat/completions 500   dial tcp 127.0.0.1:38831: connect: connection refused

127.0.0.1:45301 had been dead since 21:54:12, exactly 15m before the reap. 127.0.0.1:38831 was 30 seconds
into a healthy load and is what actually died. The client gets a 500 with nothing in the log connecting it to the
address that was reaped.

Additional context

Two things make this hard to spot from the logs:

  • The "Address is idle for too long, killing it" warning logs only the address. It never names the model it
    is about to stop, nor the live address that will actually die.
  • "[watchdog] error shutting down model: model not found" is the same bug, firing on a stale entry whose model
    happens not to be loaded at that moment. It reads like a harmless line and is the clearest early tell that stale
    entries are accumulating.

There is also a secondary effect: untrack does delete(wd.modelSizes, modelID) for the model resolved from the
stale address, so reaping a dead address also discards the live model's registered size, quietly degrading
size-aware eviction.

Notes:

  • A longer idle timeout is not a fix. The duration is not the fault; any timeout arms the same kill, only later.
  • Disabling the idle watchdog avoids it, at the cost of the keep-alive behaviour it exists to provide.
  • checkBusy reads the same addressModelMap and is worth checking for the same stale-entry exposure.

Suggested fix: untrack the address when a backend stops for any reason, not only when the watchdog stopped it,
e.g. have deleteProcess notify the watchdog (an exported Untrack(address), or a ModelUnloadHook), using the
address recorded for that model at load time. Belt and braces: in checkIdle, before shutting down, confirm the
address being reaped is still the one associated with the currently-loaded process for that model, and skip the
shutdown if it is not.

Lingua principale
Go
Stelle
49.2k
Fork
4.5k
Merge medio
1g 7h
PR unite (30g)
357

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di mudler/LocalAI

Tutte le issue di mudler/LocalAI

Issue simili

Altre issue su Go

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.