Adaptive scaling and dask-jobqueue goes into endless loop when a job launches several worker processes (was: Different configs result in worker death)
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 42/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- python
- Ambito
- distributed-systems, hpc
Direzione di ricerca
Riproduci il comportamento segnalato con SLURMCluster usando le configurazioni threaded, process-only e balanced, quindi confronta cluster.adapt con cluster.scale. Inizia dai punti di ingresso del ridimensionamento adattivo e del ciclo di vita dei worker; il lavoro è completo quando il ridimensionamento adattivo non entra più in un ciclo infinito e non perde worker mentre il carico di lavoro progredisce.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
What happened:
(Reposting from SO)
I'm using Dask Jobequeue on a Slurm supercomputer (I'll note that this is also a Cray machine). My workload includes a mix of threaded (i.e. numpy) and python workloads, so I think a balance of threads and processes would be best for my deployment (which is the default behaviour). However, in order for my jobs to run I need to use this basic configuration:
cluster = SLURMCluster(cores=20,
processes=1,
memory="60GB",
walltime='12:00:00',
...
)
cluster.adapt(minimum=0, maximum=20)
client = Client(cluster)
which is entirely threaded. The tasks also seem to take longer than I would naively expect (a large part of this is a lot of file reading/writing). Switching to purely processes, i.e.
cluster = SLURMCluster(cores=20,
processes=20,
memory="60GB",
walltime='12:00:00',
...
)
results in slurm jobs which are immediately killed by Slurm as they are launched, with the only output like:
slurmstepd: error: *** JOB 11116133 ON nid00201 CANCELLED AT 2021-04-29T17:23:25 ***
Choosing a balanced configuration (i.e. default)
cluster = SLURMCluster(cores=20,
memory="60GB",
walltime='12:00:00',
...
)
results in a strange intermediate behaviour. The task will run near to completion (i.e. 900/1000 work tasks) then a number of the workers will be killed, and the progress will drop back down to, say, 400/1000 tasks.
Further, I've found that using cluster.scale, rather than cluster.adapt, results in a successful run of the work. Perhaps the issue here is how adapt is trying to scale the number of jobs?
What you expected to happen:
I would expect that changing the balance of processes / threads shouldn't change the lifetime of a worker.
Anything else we need to know?:
Possibly related to #20 and #363
As an aside, the current configuration of processes / threads confusing, and seems to conflict with how e.g. a LocalCluster is specified. Is there any progress on #231?
Environment:
- Dask version: 2021.4.1
- Python version: 3.8.8
- Operating System: SUSE Linux Enterprise Server 12 SP3
- Install method (conda, pip, source): conda
- Lingua principale
- Python
- Stelle
- 256
- Fork
- 149
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di dask/dask-jobqueue
-
bug LSF
Difficoltà 3/5 1-2 giorni Idoneità per principianti 65/100
dask/dask-jobqueue#703 · 1 commento ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
dask/dask-jobqueue#701 ·
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 45/100
dask/dask-jobqueue#699 · 2 commenti ·
-
bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 38/100
dask/dask-jobqueue#692 · 1 commento ·
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
dask/dask-jobqueue#691 · 7 commenti ·
Tutte le issue di dask/dask-jobqueue
Issue simili
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedApertaworkflow
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
I maintainer di solito rispondono entro 1 giorno
-
New Submission: TropWATERApertametadata submission
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
canonical/content-cache-operator#163 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[submission]Apertasubmission
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
leanprover/lean-eval-submissions#1852 ·
I maintainer di solito rispondono entro 1 giorno