Adaptive scaling and dask-jobqueue goes into endless loop when a job launches several worker processes (was: Different configs result in worker death)
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 42/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 停滞
- 技術スタック
- python
調査の方向性
threaded、process-only、balanced の構成で SLURMCluster を使用して報告された動作を再現し、その後 cluster.adapt と cluster.scale を比較します。適応型スケーリングと worker のライフタイムのエントリポイントから開始します。作業の進行中に、適応型スケーリングが無限ループに入ったり worker を失ったりしなくなれば完了です。
索引モデルが issue の本文から書いたものです。
説明
What happened:
(Reposting from SO)
I'm using Dask Jobequeue on a Slurm supercomputer (I'll note that this is also a Cray machine). My workload includes a mix of threaded (i.e. numpy) and python workloads, so I think a balance of threads and processes would be best for my deployment (which is the default behaviour). However, in order for my jobs to run I need to use this basic configuration:
cluster = SLURMCluster(cores=20,
processes=1,
memory="60GB",
walltime='12:00:00',
...
)
cluster.adapt(minimum=0, maximum=20)
client = Client(cluster)
which is entirely threaded. The tasks also seem to take longer than I would naively expect (a large part of this is a lot of file reading/writing). Switching to purely processes, i.e.
cluster = SLURMCluster(cores=20,
processes=20,
memory="60GB",
walltime='12:00:00',
...
)
results in slurm jobs which are immediately killed by Slurm as they are launched, with the only output like:
slurmstepd: error: *** JOB 11116133 ON nid00201 CANCELLED AT 2021-04-29T17:23:25 ***
Choosing a balanced configuration (i.e. default)
cluster = SLURMCluster(cores=20,
memory="60GB",
walltime='12:00:00',
...
)
results in a strange intermediate behaviour. The task will run near to completion (i.e. 900/1000 work tasks) then a number of the workers will be killed, and the progress will drop back down to, say, 400/1000 tasks.
Further, I've found that using cluster.scale, rather than cluster.adapt, results in a successful run of the work. Perhaps the issue here is how adapt is trying to scale the number of jobs?
What you expected to happen:
I would expect that changing the balance of processes / threads shouldn't change the lifetime of a worker.
Anything else we need to know?:
Possibly related to #20 and #363
As an aside, the current configuration of processes / threads confusing, and seems to conflict with how e.g. a LocalCluster is specified. Is there any progress on #231?
Environment:
- Dask version: 2021.4.1
- Python version: 3.8.8
- Operating System: SUSE Linux Enterprise Server 12 SP3
- Install method (conda, pip, source): conda
- 主要言語
- Python
- スター
- 256
- フォーク
- 149
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
dask/dask-jobqueue のほかの issue
-
bug LSF
難易度 3/5 1〜2日 初心者へのやさしさ 65/100
dask/dask-jobqueue#703 · コメント 1 件 ·
-
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dask/dask-jobqueue#701 ·
-
難易度 3/5 1〜2日 初心者へのやさしさ 45/100
dask/dask-jobqueue#699 · コメント 2 件 ·
-
bug
難易度 3/5 1〜2日 初心者へのやさしさ 38/100
dask/dask-jobqueue#692 · コメント 1 件 ·
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dask/dask-jobqueue#691 · コメント 7 件 ·
dask/dask-jobqueue の issue をすべて見る
似ている issue
-
ACK_WAITING HELP_WANTED UPDATE_CS
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
OWASP/CheatSheetSeries#2458 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
BasedHardware/omi#19711 ·
メンテナーはふだん 1 日以内に返信
-
Qwen3_5MoeModel no longer returns router_logits, breaking aux loss with output_router_logits=Trueオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
huggingface/transformers#49172 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
vllm-project/vllm-metal#885 ·
メンテナーはふだん 1 日以内に返信