Multi-file install: a single part error deletes the whole install tmpdir, including completed parts and resumable partials
メンテナーはふだん 5 日以内に返信
@lstein がすでに取り組んでいます。
2026年8月24日 から。
評価
この issue はまだ評価されていません。
説明
Summary
When any single part of a multi-file install errors, _download_error_callback deletes the entire install tmpdir (model_install_default.py:1482-1489 → _safe_rmtree(install_job._install_tmpdir)) — including parts that finished completely and partials holding gigabytes of resumable progress. One transient failure (a 5xx on one file, a rename race) throws away everything the resume machinery exists to protect.
Found during the adversarial review of #9432; the rmtree behavior predates that PR, but #9432 added a new way to trip it (the sidecar rename race below).
Mechanism
Worker thread: any exception other than DownloadJobCancelledException in _do_download marks the part ERROR (download_default.py:327-330). For a part belonging to an install, _download_error_callback then pops the install job, sets it errored, cancels the multifile job, and rmtrees the whole _install_tmpdir.
Two concrete triggers:
- Transient server error on one file. A single
HTTP 500on file k of an n-file install (raised atdownload_default.py:456-459) deletes the completed files 1..k-1 and all partial progress. The next attempt starts the entire install from zero. - Sidecar rename race (new surface from #9432). The 416-promotion path stats the sidecar (
download_default.py:352) and renames it (download_default.py:446) a full network round-trip apart, with no existence guard. If the sidecar vanishes in between — user-triggeredrestart_file()runsclear_partials=True(model_install_default.py:665,:1314) against an in-flight request; manual deletion; AV/indexer interference on Windows — the rename raisesFileNotFoundError→ part ERROR → whole tmpdir deleted.
Reproduction (test sketch, deterministic)
Trigger 1 needs only two mounts — one good file, one 500 — through the install service, then assert the tmpdir is gone despite file 1 having completed.
Trigger 2 can be made deterministic at the download-queue level with an adapter that deletes the sidecar during the request, simulating the race:
class SidecarDeletingAdapter(TestAdapter):
def __init__(self, *args, sidecar: Path, **kwargs):
super().__init__(*args, **kwargs)
self._sidecar = sidecar
def send(self, request, **kwargs):
self._sidecar.unlink() # the race: sidecar vanishes mid-round-trip
return super().send(request, **kwargs)
def test_sidecar_vanishing_during_416_roundtrip(tmp_path: Path) -> None:
source = AnyHttpUrl("https://test.com/race.safetensors")
content = b"complete"
destination = tmp_path / "race.safetensors"
sidecar = destination.with_name(destination.name + ".downloading")
sidecar.write_bytes(content)
session = TestSession()
session.mount(
str(source),
SidecarDeletingAdapter(
b"", status=416, headers={"Content-Range": f"bytes */{len(content)}"}, sidecar=sidecar
),
)
queue = DownloadQueueService(requests_session=session)
queue.start()
try:
job = queue.download(source=source, dest=destination)
queue.join()
finally:
queue.stop()
# Today: FileNotFoundError escapes the rename -> job ERRORs; in an install context
# _download_error_callback then rmtrees the entire tmpdir.
assert job.status == DownloadJobStatus.ERROR
Suggested fix
- In
_download_error_callback, preserve the tmpdir when any part has resumable progress (a.downloadingfile on disk) or has already completed: write the install marker with an errored/paused status instead of rmtree, so the existing restore/restart_failedmachinery can pick it up. Only rmtree when nothing on disk is worth keeping. - Guard the 416 promotion rename: if the sidecar is missing at rename time, fall through to the pause/restart path instead of letting
FileNotFoundErrorescalate to a part ERROR.
- 主要言語
- Python
- スター
- 28.3k
- フォーク
- 3k
- 平均マージ
- 6日 22時間
- マージ済み PR(30日)
- 10
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
invoke-ai/InvokeAI のほかの issue
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
invoke-ai/InvokeAI#9608 · コメント 2 件 ·
メンテナーはふだん 5 日以内に返信
-
[bug]: Align Graph.add_edge and validate_self collector type validation対応中かも @JPPhoto が 4 日前に担当しました。 オープンbug
invoke-ai/InvokeAI#9610 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
-
[bug]: Nested Iterate execution mixes values across outer iterations対応中かも @JPPhoto が 4 日前に担当しました。 オープンbug
invoke-ai/InvokeAI#9609 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
メンテナーはふだん 5 日以内に返信
-
[bug]: 6.14.1 regression with `pytorch_cuda_alloc_conf: backend:cudaMallocAsync` — Z-Image bf16 + LoRA takes ~30 min whenever the transformer is (re)loaded (VRAM overflows into Windows shared memory)対応中かも @lstein が 6 日前に担当しました。 オープン
invoke-ai/InvokeAI#9597 · コメント 2 件 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
invoke-ai/InvokeAI の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
kornia/kornia#5263 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
approved correction metadata
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
acl-org/acl-anthology#10133 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
BasedHardware/omi#20084 ·
メンテナーはふだん 1 日以内に返信
-
bug needs-acceptance wg/evaluation-quality
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
vllm-project/semantic-router#4424 ·
メンテナーはふだん 1 日以内に返信