Checkpoint loading can mix model weights and training state from different epochs
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Đình trệ
- Lĩnh vực
- machine-learning
Hướng nghiên cứu
Start at the load_checkpoint and save_checkpoint entry points and run the minimal reproduction to observe the mismatched epoch, model weight, and optimizer state. Ensure loading chooses one checkpoint index for all requested models and training state, fails before mutation when a required file is missing, and preserves model-only directory behavior; verify the single-process and distributed cases described in the issue.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Version and installation
Source checkout: main at 94dbdf829d1a4e93e3f31ecb77713392c471388e. Also reproduced in the original #2010 implementation at 01757c816713892c38455d489376494ee7ee11e5. Python 3.13.8, PyTorch 2.12.0+cu130, Linux.
Description
load_checkpoint(..., epoch=None) independently selects the latest surviving file for the training state and for each requested model. If the newest model weights are deleted but an older weights file remains, loading succeeds with old model weights and newer optimizer/scheduler state. It reports the newer epoch, so training silently resumes from an inconsistent combination of states.
The same mismatch can occur in the other direction: a newer model file with no matching training-state file is combined with older training state.
Minimal reproduction
from pathlib import Path
from tempfile import TemporaryDirectory
import torch
from physicsnemo.utils import load_checkpoint, save_checkpoint
with TemporaryDirectory() as directory:
path = Path(directory)
model = torch.nn.Linear(1, 1, bias=False)
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
for epoch in (1, 2):
with torch.no_grad():
model.weight.fill_(epoch)
optimizer.param_groups[0]["lr"] = epoch * 0.01
save_checkpoint(path, models=model, optimizer=optimizer, epoch=epoch)
(path / "Linear.0.2.pt").unlink()
fresh = torch.nn.Linear(1, 1, bias=False)
fresh_optimizer = torch.optim.Adam(fresh.parameters(), lr=0.5)
epoch = load_checkpoint(path, models=fresh, optimizer=fresh_optimizer)
print(epoch, fresh.weight.item(), fresh_optimizer.param_groups[0]["lr"])
Observed output:
2 1.0 0.02
The returned epoch and optimizer learning rate come from epoch 2; the model weight comes from epoch 1. No exception is raised.
Expected behavior
Select one training checkpoint index and require every requested model's weights at that same index. If any required file is missing, fail clearly before changing model or training state. The caller can explicitly select an older complete checkpoint. Do not independently fall back to older or newer model files.
This should also work for automatically numbered saves, where the filename index may exist without an epoch key in the training-state payload. Preserve current behavior for directories containing only model weights.
Scope and verification
Reproduced with single-process loading and on both ranks of a two-process CPU/Gloo DTensor run, including optimizer and scheduler restoration. The filename selection is independent of the device backend. The issue exists on main and was identified while reviewing #2010; that PR will address it.
- Ngôn ngữ chính
- Python
- Star
- 3.3k
- Fork
- 802
- Merge trung bình
- 3 ngày 4 giờ
- Pull request đã merge (30 ngày)
- 28
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/physicsnemo
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
NVIDIA/physicsnemo#2036 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
NVIDIA/physicsnemo#2035 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
? - Needs Triage bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
NVIDIA/physicsnemo#2021 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
? - Needs Triage bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
NVIDIA/physicsnemo#2020 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
NVIDIA/physicsnemo#2024 · 2 bình luận · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của NVIDIA/physicsnemo
Issue tương tự
-
repo-audit
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
scverse/repo-health#20 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
phasespace-labs/palinode#232 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
collective/icalendar#1858 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Maintainer thường phản hồi trong vòng 1 ngày
-
lfx-mcp cannot supply global variables: LangflowClient drops X-LANGFLOW-GLOBAL-VAR-* from envĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
langflow-ai/langflow#15496 ·
Maintainer thường phản hồi trong vòng 1 ngày