Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Checkpoint loading can mix model weights and training state from different epochs

Aperta
#2,012 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
35/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Ferma
Stack tecnologico
python, pytorch

Direzione di ricerca

Start at the load_checkpoint and save_checkpoint entry points and run the minimal reproduction to observe the mismatched epoch, model weight, and optimizer state. Ensure loading chooses one checkpoint index for all requested models and training state, fails before mutation when a required file is missing, and preserves model-only directory behavior; verify the single-process and distributed cases described in the issue.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug

Version and installation

Source checkout: main at 94dbdf829d1a4e93e3f31ecb77713392c471388e. Also reproduced in the original #2010 implementation at 01757c816713892c38455d489376494ee7ee11e5. Python 3.13.8, PyTorch 2.12.0+cu130, Linux.

Description

load_checkpoint(..., epoch=None) independently selects the latest surviving file for the training state and for each requested model. If the newest model weights are deleted but an older weights file remains, loading succeeds with old model weights and newer optimizer/scheduler state. It reports the newer epoch, so training silently resumes from an inconsistent combination of states.

The same mismatch can occur in the other direction: a newer model file with no matching training-state file is combined with older training state.

Minimal reproduction

from pathlib import Path
from tempfile import TemporaryDirectory

import torch
from physicsnemo.utils import load_checkpoint, save_checkpoint

with TemporaryDirectory() as directory:
    path = Path(directory)
    model = torch.nn.Linear(1, 1, bias=False)
    optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
    for epoch in (1, 2):
        with torch.no_grad():
            model.weight.fill_(epoch)
        optimizer.param_groups[0]["lr"] = epoch * 0.01
        save_checkpoint(path, models=model, optimizer=optimizer, epoch=epoch)

    (path / "Linear.0.2.pt").unlink()
    fresh = torch.nn.Linear(1, 1, bias=False)
    fresh_optimizer = torch.optim.Adam(fresh.parameters(), lr=0.5)
    epoch = load_checkpoint(path, models=fresh, optimizer=fresh_optimizer)
    print(epoch, fresh.weight.item(), fresh_optimizer.param_groups[0]["lr"])

Observed output:

2 1.0 0.02

The returned epoch and optimizer learning rate come from epoch 2; the model weight comes from epoch 1. No exception is raised.

Expected behavior

Select one training checkpoint index and require every requested model's weights at that same index. If any required file is missing, fail clearly before changing model or training state. The caller can explicitly select an older complete checkpoint. Do not independently fall back to older or newer model files.

This should also work for automatically numbered saves, where the filename index may exist without an epoch key in the training-state payload. Preserve current behavior for directories containing only model weights.

Scope and verification

Reproduced with single-process loading and on both ranks of a two-process CPU/Gloo DTensor run, including optimizer and scheduler restoration. The filename selection is independent of the device backend. The issue exists on main and was identified while reviewing #2010; that PR will address it.

Lingua principale
Python
Stelle
3.3k
Fork
787
Merge medio
3g 4h
PR unite (30g)
28

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/physicsnemo

Tutte le issue di NVIDIA/physicsnemo

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.