Should we merge Dask HPC Runners in here?
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 20/100
- issue の種類
- 機能追加
- 明瞭さ
- 説明が足りない
- 活発さ
- 停滞
- 技術スタック
- python
調査の方向性
まず、リンクされている dask-runners のプロトタイプと dask/community#346 の議論を読み、続いて既存の dask-jobqueue プロジェクトと dask-mpi プロジェクトを比較します。dask-jobqueue、dask-mpi、dask-contrib、または統合パッケージのいずれに runners を含めるかについて maintainer の判断に至り、MPIRunner と SlurmRunner のスコープについて合意できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
For a while I've been playing around with this prototype repo which implements Dask Runners for HPC systems. I'm motivated to reduce the fragmentation and confusion around tooling in the Dask HPC community, so perhaps this new code should live here.
In https://github.com/dask/community/issues/346 I wrote up the difference between Dask Clusters and Dask Runners. The TL;DR is that a Cluster creates the scheduler and worker tasks directly, for example dask_jobqueue.SLURMCluster submits jobs to SLURM for each worker. A Dask Runner is different because it is invoked from within an existing allocation and populates that job with Dask processes. This the same as how Dask MPI works.
SlurmRunner Example
If I write a Python script and call it with srun -n 6 python myscript.py the script will be invoked by Slurm 6 times in parallel on 6 different nodes/cores on the HPC. The Dask Runner class then uses the Slurm process ID environment variable to decide what role reach process should play and uses the shared filesystem to bootstrap communications with a scheduler file.
# myscript.py
from dask.distributed import Client
from dask_hpc_runner import SlurmRunner
# When entering the SlurmRunner context manager processes will decide if they should be
# the client, schdeduler or a worker.
# Only process ID 1 executes the contents of the context manager.
# All other processes start the Dask components and then block here forever.
with SlurmRunner(scheduler_file="/path/to/shared/filesystem/scheduler-{job_id}.json") as runner:
# The runner object contains the scheduler address info and can be used to construct a client.
with Client(runner) as client:
# Wait for all the workers to be ready before continuing.
client.wait_for_workers(runner.n_workers)
# Then we can submit some work to the Dask scheduler.
assert client.submit(lambda x: x + 1, 10).result() == 11
assert client.submit(lambda x: x + 1, 20, workers=2).result() == 21
# When process ID 1 exits the SlurmRunner context manager it sends a graceful shutdown to the Dask processes.
Should this live in dask-jobqueue?
I'm at the point of trying to decide where this code should live within the Dask ecosystem. So far I have implemented MPIRunner and SlurmRunner as a proof-of-concept. It would be very straight forward to write runners for other batch systems provided it is possible to detect the process ID/rank from the environment.
I can imagine users choosing between SLURMCluster and SlurmRunner depending on their use case and how they want to deploy Dask. There are pros/cons to each deployment model, for example the cluster can adaptively scale, but the runner only requires a single job submission which will guarantee better node locality. So perhaps it makes sense for SlurmRunner to live here in dask-jobqueue and we can use documentation to help users choose the right one for them? (We can make the name casing more consistent).
The MPIRunner and SlurmRunner share a common base class, so I'm not sure if that means MPIRunner should also live here, or whether we should accept some code duplication and put it in dask-mpi?
Alternatively my prototype repo could just move to dask-contrib and become a separate project?
Or we could roll all of dask-jobqueue, dask-mpi and the new dask-hpc-runners into a single dask-hpc package? Or pull everything into dask-jobqueue?
The Dask HPC tooling is currently very fragmented and I'm keen to make things better, not worse. But I'm very keen to hear opinions from folks like @guillaumeeb @lesteve @jrbourbeau @kmpaul @mrocklin on what we should do here.
- 主要言語
- Python
- スター
- 256
- フォーク
- 149
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
dask/dask-jobqueue のほかの issue
-
bug LSF
難易度 3/5 1〜2日 初心者へのやさしさ 65/100
dask/dask-jobqueue#703 · コメント 1 件 ·
-
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dask/dask-jobqueue#701 ·
-
難易度 3/5 1〜2日 初心者へのやさしさ 45/100
dask/dask-jobqueue#699 · コメント 2 件 ·
-
bug
難易度 3/5 1〜2日 初心者へのやさしさ 38/100
dask/dask-jobqueue#692 · コメント 1 件 ·
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dask/dask-jobqueue#691 · コメント 7 件 ·
dask/dask-jobqueue の issue をすべて見る
似ている issue
-
bug server
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
sportsdataverse/sportsdataverse-py#641 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
googleapis/google-cloud-python#18532 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信