Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Idle watchdog reaps a stale address and shuts down a different, live backend

オープン
#12,331 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
70/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
go
領域
backend

調査の方向性

Start with pkg/model/watchdog.go and pkg/model/process.go, tracing WatchDog.Add, addressModelMap, checkIdle/checkBusy, and deleteProcess. Verify the fix by ensuring externally stopped addresses leave no stale watchdog state and that an expired old address cannot shut down a replacement backend; cover the behavior with regression tests around the watchdog lifecycle.

索引モデルが issue の本文から書いたものです。

説明

bug unconfirmed

Written on behalf of YourYoungerBrothersPug by Claude Code.

LocalAI version:

Observed on v4.9.0 (localai/localai:v4.9.0-gpu-nvidia-cuda-12). The code path is identical in v4.10.0 and on
master @ 590512d9: pkg/model/watchdog.go has not changed between them.

Environment, CPU architecture, OS, and Version:

  • Linux 6.6.114.1-microsoft-standard-WSL2 #1 SMP PREEMPT_DYNAMIC Mon Dec 1 20:46:23 UTC 2025 x86_64 GNU/Linux
    (uname -a inside the LocalAI container)
  • x86_64, Docker on WSL2, NVIDIA RTX 3060 (12 GB), CUDA 12
  • Single node. Not distributed mode: no workers, no federation.

Describe the bug

WatchDog is told when a backend starts but never when one stops, unless the watchdog itself stopped it.
Any backend torn down by another path (POST /backend/shutdown, ShutdownModel, a crash) leaves its address
behind in idleTime and addressModelMap permanently.

One idle timeout later checkIdle() fires on that dead address, resolves it to a model name, and calls
ShutdownModel(model). That stops whatever process is serving that model at that moment, which is a
different, healthy backend, and often one that is still loading.

Registration happens on load, in pkg/model/process.go:

ml.wd.Add(serverAddress, grpcControlProcess)
ml.wd.AddAddressModelMap(serverAddress, id)

There is no deregistration. untrack(address) is unexported and is called from four places, all in watchdog.go
and all watchdog-initiated evictions: checkIdle, checkBusy, collectEvictionsLocked, evictLRUModel.
deleteProcess (pkg/model/process.go) runs the unload hooks, calls Free(), stops the process, does
store.Delete(s) and cleanupProcessRuntime(process), and touches no watchdog state at all. Grepping the whole
pkg/model package for wd. finds only Add, AddAddressModelMap, RegisterModelSize, UpdateLastUsed,
EnforceLRULimit and EnforceGroupExclusivity: nothing that removes an address.

An address enters idleTime in finishRequestLocked (wd.idleTime[address] = now), so any backend that served
at least one request and was then stopped externally has a stale entry that will eventually come due. Then, in
checkIdle():

for address, t := range wd.idleTime {
    if time.Since(t) > wd.idletimeout {
        model, ok := wd.addressModelMap[address]
        ...
        modelsToShutdown = append(modelsToShutdown, model)
        wd.untrack(address)
    }
}
...
for _, model := range modelsToShutdown {
    if err := wd.pm.ShutdownModel(model); err != nil { ... }
}

The timer belongs to the address; the shutdown names the model. Nothing checks that the address being
reaped is still the address currently serving that model.

To Reproduce

  1. Run a single node with LOCALAI_WATCHDOG_IDLE=true. Set LOCALAI_WATCHDOG_IDLE_TIMEOUT=1m to see it within a
    minute; the 15m default behaves identically, just later.
  2. Send one request to model M, so a backend starts and an idleTime entry is created for its address.
  3. Stop it externally: POST /backend/shutdown {"model": "M"}.
  4. Send another request to M. A new backend starts on a new address. Confirm the two differ in the
    BackendLoader starting lines.
  5. Wait out the idle timeout measured from step 2, keeping M in use so the new backend never goes idle itself.

The watchdog reaps the dead address from step 2 and calls ShutdownModel("M"), stopping the healthy backend
from step 4. Any workload that stops backends deliberately arms one of these per stop, so they accumulate.

Expected behavior

An address that no longer has a running backend should not be reaped, and reaping one should never stop a
different, live process. Concretely: after step 3, the stale entry should be gone, and step 5 should be a no-op.

Logs

Observed ten times in three hours on one box, each exactly the 15m timeout after a different, already-dead
address for that model had stopped. The worst case is a kill landing during a load, because a load is the longest
thing that happens to a model. Trimmed to the relevant lines (these appear at info level; the first line is an
annotation, not log output):

          [a request for granite-4.1-8b arrives]
22:08:42  BackendLoader starting  modelID=granite-4.1-8b  ->  127.0.0.1:38831
22:09:11  [WatchDog] Address is idle for too long, killing it  address=127.0.0.1:45301
22:09:32  Backend process stopped  id=granite-4.1-8b  address=127.0.0.1:38831  exitCode=-1
22:09:32  POST /v1/chat/completions 500   dial tcp 127.0.0.1:38831: connect: connection refused

127.0.0.1:45301 had been dead since 21:54:12, exactly 15m before the reap. 127.0.0.1:38831 was 30 seconds
into a healthy load and is what actually died. The client gets a 500 with nothing in the log connecting it to the
address that was reaped.

Additional context

Two things make this hard to spot from the logs:

  • The "Address is idle for too long, killing it" warning logs only the address. It never names the model it
    is about to stop, nor the live address that will actually die.
  • "[watchdog] error shutting down model: model not found" is the same bug, firing on a stale entry whose model
    happens not to be loaded at that moment. It reads like a harmless line and is the clearest early tell that stale
    entries are accumulating.

There is also a secondary effect: untrack does delete(wd.modelSizes, modelID) for the model resolved from the
stale address, so reaping a dead address also discards the live model's registered size, quietly degrading
size-aware eviction.

Notes:

  • A longer idle timeout is not a fix. The duration is not the fault; any timeout arms the same kill, only later.
  • Disabling the idle watchdog avoids it, at the cost of the keep-alive behaviour it exists to provide.
  • checkBusy reads the same addressModelMap and is worth checking for the same stale-entry exposure.

Suggested fix: untrack the address when a backend stops for any reason, not only when the watchdog stopped it,
e.g. have deleteProcess notify the watchdog (an exported Untrack(address), or a ModelUnloadHook), using the
address recorded for that model at load time. Belt and braces: in checkIdle, before shutting down, confirm the
address being reaped is still the one associated with the currently-loaded process for that model, and skip the
shutdown if it is not.

主要言語
Go
スター
49.2k
フォーク
4.5k
平均マージ
1日 7時間
マージ済み PR(30日)
362

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

mudler/LocalAI のほかの issue

mudler/LocalAI の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。