jobs not resubmitted in `SLURMCluster` after graceful closure via `--lifetime`
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- python
- Ambito
- distributed-systems, infrastructure
Direzione di ricerca
Inizia con l’esempio SLURMCluster e con le impostazioni della durata di vita dei worker mostrate nell’issue, quindi riproduci il comportamento con --lifetime e --lifetime-stagger osservando al contempo lo scaling adattivo. Traccia il modo in cui gli arresti aggraziati dei worker che hanno raggiunto la durata di vita (worker-lifetime-reached) vengono segnalati al cluster e allo scheduler. Il lavoro è completato quando vengono inviati job SLURM sostitutivi dopo la chiusura aggraziata dei worker, con un test di regressione che copra questo comportamento.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi there,
I set up a SLURMCluster following the example below, including --lifetime and --lifetime-stagger to ensure jobs are closed gracefully by the dask scheduler.
import time
from dask.distributed import Client
from dask_jobqueue import SLURMCluster
GPU_CONFIG = {
'queue': 'batch_gpu',
'cores': 8,
'memory': '8GB',
'job_extra_directives': [
'--gres=gpu:1',
],
'walltime': '00:05:00',
'worker_extra_args': ["--lifetime", "10s", "--lifetime-stagger", "10s"],
}
cluster = SLURMCluster(**GPU_CONFIG)
client = Client(cluster)
cluster.adapt(minimum_jobs=1, maximum_jobs=5)
while True:
print(client)
time.sleep(5)
My expectation was that this would submit new jobs as the original ones were closed I am not seeing this behavior. I am fairly sure that when I used this exact setup a few years ago it worked correctly...
The stdout below clearly shows that the initial two jobs submit and connect to the cluster before being killed and not respawning.
<Client: 'tcp://10.164.24.30:39231' processes=0 threads=0, memory=0 B>
<Client: 'tcp://10.164.24.30:39231' processes=8 threads=16, memory=14.88 GiB>
<Client: 'tcp://10.164.24.30:39231' processes=7 threads=14, memory=13.02 GiB>
<Client: 'tcp://10.164.24.30:39231' processes=4 threads=8, memory=7.44 GiB>
<Client: 'tcp://10.164.24.30:39231' processes=4 threads=8, memory=7.44 GiB>
<Client: 'tcp://10.164.24.30:39231' processes=0 threads=0, memory=0 B>
<Client: 'tcp://10.164.24.30:39231' processes=0 threads=0, memory=0 B>
<Client: 'tcp://10.164.24.30:39231' processes=0 threads=0, memory=0 B>
<Client: 'tcp://10.164.24.30:39231' processes=0 threads=0, memory=0 B>
The following worker log shows the worker processes die as they reach their lifetime and close gracefully as expected.
2025-06-09 17:06:59,476 - distributed.nanny - INFO - Start Nanny at: 'tcp://10.164.25.82:35379'
2025-06-09 17:06:59,480 - distributed.nanny - INFO - Start Nanny at: 'tcp://10.164.25.82:33079'
2025-06-09 17:06:59,481 - distributed.nanny - INFO - Start Nanny at: 'tcp://10.164.25.82:41377'
2025-06-09 17:06:59,483 - distributed.nanny - INFO - Start Nanny at: 'tcp://10.164.25.82:35255'
2025-06-09 17:06:59,900 - distributed.diskutils - INFO - Found stale lock file and directory '/tmp/dask-scratch-space/worker-p715og9m', purging
2025-06-09 17:06:59,901 - distributed.diskutils - INFO - Found stale lock file and directory '/tmp/dask-scratch-space/worker-9_cdptn2', purging
2025-06-09 17:07:00,230 - distributed.worker - INFO - Start worker at: tcp://10.164.25.82:40503
2025-06-09 17:07:00,230 - distributed.worker - INFO - Listening to: tcp://10.164.25.82:40503
2025-06-09 17:07:00,230 - distributed.worker - INFO - Worker name: SLURMCluster-0-0
2025-06-09 17:07:00,230 - distributed.worker - INFO - dashboard at: 10.164.25.82:44353
2025-06-09 17:07:00,230 - distributed.worker - INFO - Waiting to connect to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,231 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,231 - distributed.worker - INFO - Threads: 2
2025-06-09 17:07:00,231 - distributed.worker - INFO - Memory: 1.86 GiB
2025-06-09 17:07:00,231 - distributed.worker - INFO - Local Directory: /tmp/dask-scratch-space/worker-guwx53y4
2025-06-09 17:07:00,231 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,240 - distributed.worker - INFO - Start worker at: tcp://10.164.25.82:45191
2025-06-09 17:07:00,240 - distributed.worker - INFO - Listening to: tcp://10.164.25.82:45191
2025-06-09 17:07:00,240 - distributed.worker - INFO - Worker name: SLURMCluster-0-3
2025-06-09 17:07:00,240 - distributed.worker - INFO - dashboard at: 10.164.25.82:46333
2025-06-09 17:07:00,240 - distributed.worker - INFO - Waiting to connect to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,240 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,240 - distributed.worker - INFO - Threads: 2
2025-06-09 17:07:00,240 - distributed.worker - INFO - Memory: 1.86 GiB
2025-06-09 17:07:00,240 - distributed.worker - INFO - Local Directory: /tmp/dask-scratch-space/worker-0j9vxv6a
2025-06-09 17:07:00,240 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,242 - distributed.worker - INFO - Start worker at: tcp://10.164.25.82:42045
2025-06-09 17:07:00,242 - distributed.worker - INFO - Listening to: tcp://10.164.25.82:42045
2025-06-09 17:07:00,242 - distributed.worker - INFO - Start worker at: tcp://10.164.25.82:36879
2025-06-09 17:07:00,242 - distributed.worker - INFO - Worker name: SLURMCluster-0-1
2025-06-09 17:07:00,242 - distributed.worker - INFO - dashboard at: 10.164.25.82:43743
2025-06-09 17:07:00,242 - distributed.worker - INFO - Listening to: tcp://10.164.25.82:36879
2025-06-09 17:07:00,242 - distributed.worker - INFO - Waiting to connect to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,242 - distributed.worker - INFO - Worker name: SLURMCluster-0-2
2025-06-09 17:07:00,242 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,242 - distributed.worker - INFO - dashboard at: 10.164.25.82:39849
2025-06-09 17:07:00,242 - distributed.worker - INFO - Waiting to connect to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,242 - distributed.worker - INFO - Threads: 2
2025-06-09 17:07:00,242 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,243 - distributed.worker - INFO - Memory: 1.86 GiB
2025-06-09 17:07:00,243 - distributed.worker - INFO - Local Directory: /tmp/dask-scratch-space/worker-n29701zb
2025-06-09 17:07:00,243 - distributed.worker - INFO - Threads: 2
2025-06-09 17:07:00,243 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,243 - distributed.worker - INFO - Memory: 1.86 GiB
2025-06-09 17:07:00,243 - distributed.worker - INFO - Local Directory: /tmp/dask-scratch-space/worker-cxvq1qyq
2025-06-09 17:07:00,243 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,245 - distributed.worker - INFO - Starting Worker plugin shuffle
2025-06-09 17:07:00,246 - distributed.worker - INFO - Registered to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,246 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,246 - distributed.core - INFO - Starting established connection to tcp://10.164.24.30:39231
2025-06-09 17:07:00,253 - distributed.worker - INFO - Starting Worker plugin shuffle
2025-06-09 17:07:00,253 - distributed.worker - INFO - Registered to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,253 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,254 - distributed.core - INFO - Starting established connection to tcp://10.164.24.30:39231
2025-06-09 17:07:00,255 - distributed.worker - INFO - Starting Worker plugin shuffle
2025-06-09 17:07:00,256 - distributed.worker - INFO - Registered to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,256 - distributed.worker - INFO - Starting Worker plugin shuffle
2025-06-09 17:07:00,256 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,256 - distributed.worker - INFO - Registered to: tcp://10.164.24.30:39231
2025-06-09 17:07:00,256 - distributed.worker - INFO - -------------------------------------------------
2025-06-09 17:07:00,256 - distributed.core - INFO - Starting established connection to tcp://10.164.24.30:39231
2025-06-09 17:07:00,256 - distributed.core - INFO - Starting established connection to tcp://10.164.24.30:39231
2025-06-09 17:07:02,099 - distributed.worker - INFO - Closing worker gracefully: tcp://10.164.25.82:45191. Reason: worker-lifetime-reached
2025-06-09 17:07:02,100 - distributed.worker - INFO - Stopping worker at tcp://10.164.25.82:45191. Reason: worker-lifetime-reached
2025-06-09 17:07:02,102 - distributed.nanny - INFO - Closing Nanny gracefully at 'tcp://10.164.25.82:33079'. Reason: worker-lifetime-reached
2025-06-09 17:07:02,102 - distributed.worker - INFO - Removing Worker plugin shuffle
2025-06-09 17:07:02,103 - distributed.core - INFO - Connection to tcp://10.164.24.30:39231 has been closed.
2025-06-09 17:07:02,103 - distributed.nanny - INFO - Worker closed
2025-06-09 17:07:04,104 - distributed.nanny - ERROR - Worker process died unexpectedly
2025-06-09 17:07:04,202 - distributed.nanny - INFO - Closing Nanny at 'tcp://10.164.25.82:33079'. Reason: nanny-close-gracefully
2025-06-09 17:07:04,203 - distributed.nanny - INFO - Nanny at 'tcp://10.164.25.82:33079' closed.
2025-06-09 17:07:04,590 - distributed.worker - INFO - Closing worker gracefully: tcp://10.164.25.82:36879. Reason: worker-lifetime-reached
2025-06-09 17:07:04,592 - distributed.worker - INFO - Stopping worker at tcp://10.164.25.82:36879. Reason: worker-lifetime-reached
2025-06-09 17:07:04,593 - distributed.nanny - INFO - Closing Nanny gracefully at 'tcp://10.164.25.82:35255'. Reason: worker-lifetime-reached
2025-06-09 17:07:04,593 - distributed.worker - INFO - Removing Worker plugin shuffle
2025-06-09 17:07:04,594 - distributed.core - INFO - Connection to tcp://10.164.24.30:39231 has been closed.
2025-06-09 17:07:04,594 - distributed.nanny - INFO - Worker closed
2025-06-09 17:07:06,693 - distributed.nanny - INFO - Closing Nanny at 'tcp://10.164.25.82:35255'. Reason: nanny-close-gracefully
2025-06-09 17:07:06,693 - distributed.nanny - INFO - Nanny at 'tcp://10.164.25.82:35255' closed.
2025-06-09 17:07:15,959 - distributed.worker - INFO - Closing worker gracefully: tcp://10.164.25.82:42045. Reason: worker-lifetime-reached
2025-06-09 17:07:15,961 - distributed.worker - INFO - Stopping worker at tcp://10.164.25.82:42045. Reason: worker-lifetime-reached
2025-06-09 17:07:15,962 - distributed.nanny - INFO - Closing Nanny gracefully at 'tcp://10.164.25.82:35379'. Reason: worker-lifetime-reached
2025-06-09 17:07:15,962 - distributed.worker - INFO - Removing Worker plugin shuffle
2025-06-09 17:07:15,963 - distributed.core - INFO - Connection to tcp://10.164.24.30:39231 has been closed.
2025-06-09 17:07:15,964 - distributed.nanny - INFO - Worker closed
2025-06-09 17:07:16,803 - distributed.worker - INFO - Closing worker gracefully: tcp://10.164.25.82:40503. Reason: worker-lifetime-reached
2025-06-09 17:07:16,805 - distributed.worker - INFO - Stopping worker at tcp://10.164.25.82:40503. Reason: worker-lifetime-reached
2025-06-09 17:07:16,806 - distributed.nanny - INFO - Closing Nanny gracefully at 'tcp://10.164.25.82:41377'. Reason: worker-lifetime-reached
2025-06-09 17:07:16,806 - distributed.worker - INFO - Removing Worker plugin shuffle
2025-06-09 17:07:16,807 - distributed.core - INFO - Connection to tcp://10.164.24.30:39231 has been closed.
2025-06-09 17:07:16,808 - distributed.nanny - INFO - Worker closed
2025-06-09 17:07:17,965 - distributed.nanny - ERROR - Worker process died unexpectedly
2025-06-09 17:07:18,063 - distributed.nanny - INFO - Closing Nanny at 'tcp://10.164.25.82:35379'. Reason: nanny-close-gracefully
2025-06-09 17:07:18,063 - distributed.nanny - INFO - Nanny at 'tcp://10.164.25.82:35379' closed.
2025-06-09 17:07:18,993 - distributed.nanny - INFO - Closing Nanny at 'tcp://10.164.25.82:41377'. Reason: nanny-close-gracefully
2025-06-09 17:07:18,993 - distributed.nanny - INFO - Nanny at 'tcp://10.164.25.82:41377' closed.
2025-06-09 17:07:18,993 - distributed.dask_worker - INFO - End worker
I also confirmed that I see no new workers in the SLURM job queue with squeue.
Is this a bug or am I doing something obviously wrong here? Many thanks in advance for any help with this
- Lingua principale
- Python
- Stelle
- 256
- Fork
- 149
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di dask/dask-jobqueue
-
bug LSF
Difficoltà 3/5 1-2 giorni Idoneità per principianti 65/100
dask/dask-jobqueue#703 · 1 commento ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
dask/dask-jobqueue#701 ·
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 45/100
dask/dask-jobqueue#699 · 2 commenti ·
-
bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 38/100
dask/dask-jobqueue#692 · 1 commento ·
-
Difficoltà 2/5 Mezza giornata Idoneità per principianti 45/100
dask/dask-jobqueue#686 · 3 commenti ·
Tutte le issue di dask/dask-jobqueue
Issue simili
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedApertaworkflow
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
I maintainer di solito rispondono entro 1 giorno
-
New Submission: TropWATERApertametadata submission
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
canonical/content-cache-operator#163 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[submission]Apertasubmission
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
leanprover/lean-eval-submissions#1852 ·
I maintainer di solito rispondono entro 1 giorno