Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

[Bug]: engine:status re-runs the readiness probe under the engine lock on every call; a slow engine /health starves the model-list sweep and, polled ~1/s, leaked a TP=4 head to death

未關閉
#84 2 則留言 0 個 reaction 已指派 1 人 在 GitHub 檢視

維護者通常 5 天內回覆

@kjlubick 已經在處理了。

開始於 2026年9月18日。

評估

這個 Issue 還沒有評估資料。

描述

Area

Engine or model management

User problem

Every engine:status call re-runs the engine's readiness probe while holding that engine's lifecycle lock, and several PAIR components poll status independently. On an engine whose readiness endpoint is cheap that is invisible. On an engine whose /health does real work it has two consequences we hit in production this week on a 4-node tensor-parallel SGLang head (DGX Spark, GB10):

  1. Model-list starvation. The broker's advertiser (every 5 s), the loaded-model watcher (every 5 s) and the desktop's remote status poll (every 10 s) each call engine:status. StatusAtPort takes st.opMu, then reconcilePresence runs probe(ready) and probe(identity) with no caching. SGLang's /health performs a short generation and takes ~1.0 s on this build, so the mutex was held essentially 100% of the time and ModelsResult's sweep (which needs Status first) never got in. GET :14322/v1/models on that node hung for 40 s+ indefinitely; peers piled up hundreds of CLOSE-WAIT sockets; the desktop logged remote engine status ... unavailable every 10 s. A standalone engine-manager with no broker traffic answered in 16 ms. Restarting engine-manager did not help.
  2. The probe load itself leaked memory. The head's container log shows 72,144 GET /health and 69,817 GET /get_model_info over a 20 h run, ~1/s each, with zero user requests for the final 30 min. The head's MemAvailable declined monotonically from 8.7 GB (00:20) to 2.5 GB (16:20) while the three worker ranks stayed flat, then earlyoom SIGTERMed the scheduler at 16:28 and the TP group died. After the crash the head returned to its idle baseline, so the growth was inside the front-end processes only rank 0 runs. Pointing the probes at /get_model_info (~1 ms, no generation) dropped /health traffic from ~3,500/h to the container's own healthcheck and the model list answers in 13 ms.

SGLang itself is not in develop yet (it lives in #50 and in my fork), but the mechanism is upstream code and applies to any engine whose readiness endpoint is not free; llama.cpp's /health under load and /v1/models on busy servers are candidates.

Where
  • services/nvpair-engine-manager/status.go: StatusAtPort → st.opMu.Lock() → reconcilePresence(context.Background(), ...) → e.probe(ctx, ready, port) then e.probe(ctx, identity, port) on every call.
  • services/nvpair-engine-manager/models.go: ModelsResult calls e.Status(name) per engine before the 5 s action budget starts; the lock wait is unbounded.
  • Pollers: nvpair-ui-broker/advertiser.go (autoAdvertiseInterval = 5 * time.Second), nvpair-engine-manager/loadedwatch.go (defaultLoadedPollSeconds = 5), the desktop's remote-get-installed loop.
Proposed fix
  • Do not re-run the readiness probe for an engine that is already adopted and healthy with a live health loop; trust the health loop's last result, or cache presence for a few seconds.
  • Do not take opMu for the read-only status path; snapshot state, probe outside the lock.
  • Treat the manifest's identity endpoint as the default readiness/health probe and require an explicit opt-in for anything that generates.
Workaround for operators

A per-engine manifest override (engines/sglang.json) pointing runtime.ready.http and runtime.health.http at /get_model_info, plus the advertiser change in https://github.com/jlacroix82/Personal-AI-Router/commit/c5b9be7 (on feat/vllm-sglang).

Environment

PAIR 0.1.1 services (engine-manager 0.21.0 / broker 0.42.2 as built from feat/vllm-sglang at ff26f5b), Linux arm64, DGX Spark x4 per TP group, SGLang lmsysorg/sglang:dev-dsv41 serving DeepSeek-V4.1-Flash. Related: #37 (probe connection reuse), #50 (SGLang engine), #24 (external backends).

主要語言
Go
星號
1.6k
分支
266
平均合併
3 天 15 小時
30 天內合併 PR
23

環境準備

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

NVIDIA/Personal-AI-Router 的其他 Issue

查看 NVIDIA/Personal-AI-Router 的全部 Issue

相似的 Issue

更多 Go Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。