Lightning-AI/pytorch-lightning

MLFlow Logger Makes a New Run When Resuming from hpc Checkpoint

开放

#6,205 创建于 2021年2月25日

 (8 条评论) (0 个反应) (0 位负责人)Python (3,233 个派生)batch import
bugcheckpointingenvironment: slurmhelp wantedlogger: mlflowpriority: 2won't fix

仓库指标

星标
 (26,687 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

🐛 Bug

Currently the MLFlowLogger creates a new run when resuming from an hpc checkpoint, e.g., after preemption by slurm and requeuing. Runs are an MLFlow concept that groups things in their UI, so when resuming after requeue, it should really be reusing the run ID. I think this can be patched into the hpc checkpoint using the logger which I believe exposes the run ID. This can also be seen on the v_num on the progress bar which changes after preemption (in general that v_num probably shouldnt be changing in this case). I'm happy to attempt to PR this if the owners agree that it's a bug.

To Reproduce

Use MLFlowLogger on a slurm cluster and watch the mlflow UI when preemption happens, there will be a new run created.

Expected behavior

Runs are grouped neatly on the MLFlow UI

Environment

  • CUDA: - GPU: - available: False - version: 10.2
  • Packages: - numpy: 1.20.1 - pyTorch_debug: False - pyTorch_version: 1.7.1 - pytorch-lightning: 1.2.0 - tqdm: 4.57.0
  • System: - OS: Linux - architecture: - 64bit - ELF - processor: x86_64 - python: 3.8.1 - version: #1 SMP Thu Jan 21 16:15:07 EST 2021

cc @awaelchli @ananthsub @ninginthecloud @rohitgr7 @tchaton @akihironitta

贡献者指南