Multi-file install: a single part error deletes the whole install tmpdir, including completed parts and resumable partials
Los mantenedores suelen responder en 5 días
@lstein ya está trabajando en esto.
Desde el 24/8/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Summary
When any single part of a multi-file install errors, _download_error_callback deletes the entire install tmpdir (model_install_default.py:1482-1489 → _safe_rmtree(install_job._install_tmpdir)) — including parts that finished completely and partials holding gigabytes of resumable progress. One transient failure (a 5xx on one file, a rename race) throws away everything the resume machinery exists to protect.
Found during the adversarial review of #9432; the rmtree behavior predates that PR, but #9432 added a new way to trip it (the sidecar rename race below).
Mechanism
Worker thread: any exception other than DownloadJobCancelledException in _do_download marks the part ERROR (download_default.py:327-330). For a part belonging to an install, _download_error_callback then pops the install job, sets it errored, cancels the multifile job, and rmtrees the whole _install_tmpdir.
Two concrete triggers:
- Transient server error on one file. A single
HTTP 500on file k of an n-file install (raised atdownload_default.py:456-459) deletes the completed files 1..k-1 and all partial progress. The next attempt starts the entire install from zero. - Sidecar rename race (new surface from #9432). The 416-promotion path stats the sidecar (
download_default.py:352) and renames it (download_default.py:446) a full network round-trip apart, with no existence guard. If the sidecar vanishes in between — user-triggeredrestart_file()runsclear_partials=True(model_install_default.py:665,:1314) against an in-flight request; manual deletion; AV/indexer interference on Windows — the rename raisesFileNotFoundError→ part ERROR → whole tmpdir deleted.
Reproduction (test sketch, deterministic)
Trigger 1 needs only two mounts — one good file, one 500 — through the install service, then assert the tmpdir is gone despite file 1 having completed.
Trigger 2 can be made deterministic at the download-queue level with an adapter that deletes the sidecar during the request, simulating the race:
class SidecarDeletingAdapter(TestAdapter):
def __init__(self, *args, sidecar: Path, **kwargs):
super().__init__(*args, **kwargs)
self._sidecar = sidecar
def send(self, request, **kwargs):
self._sidecar.unlink() # the race: sidecar vanishes mid-round-trip
return super().send(request, **kwargs)
def test_sidecar_vanishing_during_416_roundtrip(tmp_path: Path) -> None:
source = AnyHttpUrl("https://test.com/race.safetensors")
content = b"complete"
destination = tmp_path / "race.safetensors"
sidecar = destination.with_name(destination.name + ".downloading")
sidecar.write_bytes(content)
session = TestSession()
session.mount(
str(source),
SidecarDeletingAdapter(
b"", status=416, headers={"Content-Range": f"bytes */{len(content)}"}, sidecar=sidecar
),
)
queue = DownloadQueueService(requests_session=session)
queue.start()
try:
job = queue.download(source=source, dest=destination)
queue.join()
finally:
queue.stop()
# Today: FileNotFoundError escapes the rename -> job ERRORs; in an install context
# _download_error_callback then rmtrees the entire tmpdir.
assert job.status == DownloadJobStatus.ERROR
Suggested fix
- In
_download_error_callback, preserve the tmpdir when any part has resumable progress (a.downloadingfile on disk) or has already completed: write the install marker with an errored/paused status instead of rmtree, so the existing restore/restart_failedmachinery can pick it up. Only rmtree when nothing on disk is worth keeping. - Guard the 416 promotion rename: if the sidecar is missing at rename time, fall through to the pause/restart path instead of letting
FileNotFoundErrorescalate to a part ERROR.
- Lenguaje dominante
- Python
- Estrellas
- 28.3k
- Forks
- 3k
- Merge medio
- 6 d 22 h
- PR fusionados (30 d)
- 10
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de invoke-ai/InvokeAI
-
[enhancement]: UpscalingAbiertoenhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
invoke-ai/InvokeAI#9608 · 2 comentarios ·
Los mantenedores suelen responder en 5 días
-
[bug]: Align Graph.add_edge and validate_self collector type validationPosiblemente ocupada @JPPhoto la tomó hace 4 días. Abiertobug
invoke-ai/InvokeAI#9610 · 1 asignado ·
Los mantenedores suelen responder en 5 días
-
[bug]: Nested Iterate execution mixes values across outer iterationsPosiblemente ocupada @JPPhoto la tomó hace 4 días. Abiertobug
invoke-ai/InvokeAI#9609 · 1 asignado ·
Los mantenedores suelen responder en 5 días
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
Los mantenedores suelen responder en 5 días
-
[bug]: 6.14.1 regression with `pytorch_cuda_alloc_conf: backend:cudaMallocAsync` — Z-Image bf16 + LoRA takes ~30 min whenever the transformer is (re)loaded (VRAM overflows into Windows shared memory)Posiblemente ocupada @lstein la tomó hace 5 días. Abierto
invoke-ai/InvokeAI#9597 · 2 comentarios · 1 asignado ·
Los mantenedores suelen responder en 5 días
Todos los issues de invoke-ai/InvokeAI
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
letsencrypt/cp-cps#353 ·
-
Marble Madness II is missingAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
DOI-USGS/pywatershed#421 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
python-pillow/Pillow#10087 · 1 comentario ·
Los mantenedores suelen responder en 1 día