[Bug]: engine:status re-runs the readiness probe under the engine lock on every call; a slow engine /health starves the model-list sweep and, polled ~1/s, leaked a TP=4 head to death
メンテナーはふだん 5 日以内に返信
@kjlubick がすでに取り組んでいます。
2026年9月18日 から。
評価
この issue はまだ評価されていません。
説明
Area
Engine or model management
User problem
Every engine:status call re-runs the engine's readiness probe while holding that engine's lifecycle lock, and several PAIR components poll status independently. On an engine whose readiness endpoint is cheap that is invisible. On an engine whose /health does real work it has two consequences we hit in production this week on a 4-node tensor-parallel SGLang head (DGX Spark, GB10):
- Model-list starvation. The broker's advertiser (every 5 s), the loaded-model watcher (every 5 s) and the desktop's remote status poll (every 10 s) each call
engine:status.StatusAtPorttakesst.opMu, thenreconcilePresencerunsprobe(ready)andprobe(identity)with no caching. SGLang's/healthperforms a short generation and takes ~1.0 s on this build, so the mutex was held essentially 100% of the time andModelsResult's sweep (which needsStatusfirst) never got in.GET :14322/v1/modelson that node hung for 40 s+ indefinitely; peers piled up hundreds of CLOSE-WAIT sockets; the desktop loggedremote engine status ... unavailableevery 10 s. A standalone engine-manager with no broker traffic answered in 16 ms. Restarting engine-manager did not help. - The probe load itself leaked memory. The head's container log shows 72,144
GET /healthand 69,817GET /get_model_infoover a 20 h run, ~1/s each, with zero user requests for the final 30 min. The head's MemAvailable declined monotonically from 8.7 GB (00:20) to 2.5 GB (16:20) while the three worker ranks stayed flat, then earlyoom SIGTERMed the scheduler at 16:28 and the TP group died. After the crash the head returned to its idle baseline, so the growth was inside the front-end processes only rank 0 runs. Pointing the probes at/get_model_info(~1 ms, no generation) dropped/healthtraffic from ~3,500/h to the container's own healthcheck and the model list answers in 13 ms.
SGLang itself is not in develop yet (it lives in #50 and in my fork), but the mechanism is upstream code and applies to any engine whose readiness endpoint is not free; llama.cpp's /health under load and /v1/models on busy servers are candidates.
Where
services/nvpair-engine-manager/status.go:StatusAtPort→st.opMu.Lock()→reconcilePresence(context.Background(), ...)→e.probe(ctx, ready, port)thene.probe(ctx, identity, port)on every call.services/nvpair-engine-manager/models.go:ModelsResultcallse.Status(name)per engine before the 5 s action budget starts; the lock wait is unbounded.- Pollers:
nvpair-ui-broker/advertiser.go(autoAdvertiseInterval = 5 * time.Second),nvpair-engine-manager/loadedwatch.go(defaultLoadedPollSeconds = 5), the desktop's remote-get-installed loop.
Proposed fix
- Do not re-run the readiness probe for an engine that is already adopted and healthy with a live health loop; trust the health loop's last result, or cache presence for a few seconds.
- Do not take
opMufor the read-only status path; snapshot state, probe outside the lock. - Treat the manifest's
identityendpoint as the default readiness/health probe and require an explicit opt-in for anything that generates.
Workaround for operators
A per-engine manifest override (engines/sglang.json) pointing runtime.ready.http and runtime.health.http at /get_model_info, plus the advertiser change in https://github.com/jlacroix82/Personal-AI-Router/commit/c5b9be7 (on feat/vllm-sglang).
Environment
PAIR 0.1.1 services (engine-manager 0.21.0 / broker 0.42.2 as built from feat/vllm-sglang at ff26f5b), Linux arm64, DGX Spark x4 per TP group, SGLang lmsysorg/sglang:dev-dsv41 serving DeepSeek-V4.1-Flash. Related: #37 (probe connection reuse), #50 (SGLang engine), #24 (external backends).
- 主要言語
- Go
- スター
- 1.6k
- フォーク
- 266
- 平均マージ
- 3日 15時間
- マージ済み PR(30日)
- 23
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/Personal-AI-Router のほかの issue
-
enhancement
難易度 4/5 1週間以上 初心者へのやさしさ 12/100
NVIDIA/Personal-AI-Router#162 ·
メンテナーはふだん 5 日以内に返信
-
enhancement
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
NVIDIA/Personal-AI-Router#154 · コメント 1 件 ·
メンテナーはふだん 5 日以内に返信
-
[Feature]: [Ollama] Alert the user about model pull failures対応中かも @ckelseynv が 1 日前に担当しました。 オープンenhancement
難易度 3/5 1〜2日 初心者へのやさしさ 56/100
NVIDIA/Personal-AI-Router#152 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
-
[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy対応中かも @ilaigold が 2 日前に担当しました。 オープン
難易度 3/5 半日 初心者へのやさしさ 84/100
NVIDIA/Personal-AI-Router#146 ·
メンテナーはふだん 5 日以内に返信
-
Make model download cancellation asynchronous対応中かも @ckelseynv が 18 日前に担当しました。 オープンenhancement
NVIDIA/Personal-AI-Router#115 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
NVIDIA/Personal-AI-Router の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 75/100
slavakurilyak/awesome-ai-agents#742 ·
メンテナーはふだん 1 日以内に返信
-
bug go
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
genkit-ai/genkit#6761 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 87/100
メンテナーはふだん 2 日以内に返信