Derive the model download and cancellation timeouts from measurement
メンテナーはふだん 5 日以内に返信
@ckelseynv がすでに取り組んでいます。
2026年9月23日 から。
評価
この issue はまだ評価されていません。
説明
Model download and cancellation are covered by a chain of timeouts in the desktop app, the broker, the engine manager, and the terminal interface. Reviewing #111 turned up several that are not wrong so much as unjustified: they were picked independently, and a few cannot fire or fire far too early. None of them produces a wrong user-visible outcome today — the ones that did were fixed in that pull request — so this is hardening, not a bug report.
The common thread is that each budget was chosen on its own rather than derived from the thing it waits on. The goal is to make them derived, and to measure the real durations first so the derivation rests on evidence.
Budgets that cannot apply, or apply too soon
PULL_TIMEOUT_MSis unreachable. The desktop allows six hours forpull_model, but the engine manager'sactionTimeoutis 30 minutes and compile-time only, so it cuts every pull first. The real download ceiling is 30 minutes, and the desktop comment describing "a multi-GB download on a slow link" describes a budget that can never take effect. Decide deliberately whetherpull_modelis exempt from the action ceiling or the value becomes configurable.- Post-pull
list_modelsgets 10s by omission. It inherits thejson-rpc-subprocess.tsdefault while the dispatch it invokes is budgeted at 30 minutes. For LM Studio that call is a CLI shell-out. PENDING_TIMEOUT_MSis 60s. It expires the optimistic button cover one minute into a download that can legitimately run for half an hour.- The terminal interface's
callTimeoutis 35s, below the worst-case cancel. It is tolerated only because the terminal interface swallowsDeadlineExceeded. - Cold dial versus the desktop's cancel budget. A first contact with a peer pays dial and TLS handshake before the cancel's own wait begins, and the desktop budget does not account for it.
Supporting work
- Instrumentation first. Phase-duration debug logging around a cancel — interrupt to acknowledgement, acknowledgement to exit, exit to cleanup complete — so the budgets can be set from measurement rather than assumption. Operational metadata only, never prompts or response bodies.
- Derive rather than pick. Once measured, express each budget in terms of the layer beneath it so the two cannot drift apart, and add a check that the ordering still holds.
- A permanent
-racegate in CI. The race detector does not run in CI today. It was run manually for #111 in a throwaway container and found nothing, but nothing keeps it that way. - A TypeScript parity check for the budget relationships that currently live only in comments.
Smaller items deferred from the same review
lmspull.go— give LM Studio's partial-file cleanup the budgeted retry loopcleanupOllamaAfterCancelalready has. Skipping a busy file already prevents a wrong deletion; the retry only reduces how often a partial is left behind, so it is an improvement rather than a fix.pullcancel.go— compare witherrors.Israther than==. Defensive; there is no current failure case.- Progress scanning — scan only the newly appended chunk and truncate on a rune boundary. Bounded to roughly 8KiB of ASCII today, so the only effect of a non-ASCII rune is cosmetically garbled diagnostic text.
- Classifying pre-existing partials — sample them for movement before a pull starts, so a file already growing is known to belong to another client. This is a more principled attribution rule than the current pre-pull snapshot, but it is new machinery plus startup latency.
Related
Most of these budgets exist only because a cancel blocks until it finishes. See #115, on making cancellation asynchronous, which would remove the need for several of them outright.
- 主要言語
- Go
- スター
- 1.6k
- フォーク
- 266
- 平均マージ
- 3日 15時間
- マージ済み PR(30日)
- 23
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/Personal-AI-Router のほかの issue
-
enhancement
難易度 4/5 1週間以上 初心者へのやさしさ 12/100
NVIDIA/Personal-AI-Router#162 ·
メンテナーはふだん 5 日以内に返信
-
enhancement
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
NVIDIA/Personal-AI-Router#154 · コメント 1 件 ·
メンテナーはふだん 5 日以内に返信
-
[Feature]: [Ollama] Alert the user about model pull failures対応中かも @ckelseynv が 1 日前に担当しました。 オープンenhancement
難易度 3/5 1〜2日 初心者へのやさしさ 56/100
NVIDIA/Personal-AI-Router#152 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
-
[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy対応中かも @ilaigold が 2 日前に担当しました。 オープン
難易度 3/5 半日 初心者へのやさしさ 84/100
NVIDIA/Personal-AI-Router#146 ·
メンテナーはふだん 5 日以内に返信
-
Make model download cancellation asynchronous対応中かも @ckelseynv が 17 日前に担当しました。 オープンenhancement
NVIDIA/Personal-AI-Router#115 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
NVIDIA/Personal-AI-Router の issue をすべて見る
似ている issue
-
bug triage
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
FairwindsOps/nova#484 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
automated-analysis code-quality cookie
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
github/gh-aw#67517 · コメント 3 件 ·
メンテナーはふだん 1 日以内に返信
-
[otelcol] print-config help text still requires the removed otelcol.printInitialConfig feature gate対応中かも @girishkvs が今日担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
open-telemetry/opentelemetry-collector#16143 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug good first issue load-balancing
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
ktrubilo9/edge-proxy#53 ·