Resource allocation on SLURM cluster
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 停滞
- 技術スタック
- python
調査の方向性
報告された Python スクリプトと SLURMCluster 設定から始め、worker と scheduler のログを調べながら、submit、map、adapt を使って動作を再現します。要求された GPU リソース、worker プロセス、タスクの配置を比較します。完了の条件は、文書化された設定によって意図した worker とタスクが確実にスケジュールされ、アイドル状態の worker や worker の繰り返しの停止がないことです。
索引モデルが issue の本文から書いたものです。
説明
Describe the issue:
This is likely a misunderstanding of how to correctly use Dask to deploy cluster jobs. However, the terminology in the documentation suggests what should be happening, so I also see this as a type of bug as the functionality is so different to what one might expect.
I am trying to train a large number of machine learning models on a SLURM cluster. Each Node has 64 cores and 4 GPUs. I want to run each of my model with 1 GPU and 16 cores so I can, theoretically, get 4 models on each Node and maximise my resources.
My input script is summarised as follows:
def train(index: int):
"""
Run the experiment.
"""
class Network(nn.Module):
"""
Perceptron network.
"""
@nn.compact
def __call__(self, x):
"""
Call method for the network
"""
x = nn.Dense(2, use_bias=False)(x)
return nn.sigmoid(x)
generator = DecisionBoundaryGenerator(ds_size, discriminator="line", one_hot=True)
model = nl.models.FlaxModel(
flax_module=Network(),
optimizer=optax.adam(0.01),
input_shape=(1, 2)
)
# Prepare the recorders
train_recorder = nl.training_recording.JaxRecorder(
name=f"ce-perceptron/train_recorder_{index}",
loss=True,
entropy=True,
trace=True,
accuracy=True,
magnitude_variance=True,
update_rate=1,
)
test_recorder = nl.training_recording.JaxRecorder(
name=f"ce-perceptron/test_recorder_{index}", loss=True, accuracy=True, update_rate=1,
)
train_recorder.instantiate_recorder(data_set=generator.train_ds)
test_recorder.instantiate_recorder(data_set=generator.test_ds)
trainer = nl.training_strategies.SimpleTraining(
model=model,
loss_fn=nl.loss_functions.CrossEntropyLoss(),
accuracy_fn=nl.accuracy_functions.LabelAccuracy(),
recorders=[train_recorder, test_recorder],
)
_ = trainer.train_model(
train_ds=generator.train_ds,
test_ds=generator.test_ds,
batch_size=128,
epochs=5000,
)
indices = np.linspace(1, 20, 20, dtype=int)
cluster = SLURMCluster(
cores=16,
processes=1,
memory="64GB",
queue="Anonymised",
walltime="01:00:00",
death_timeout="15s",
worker_extra_args=["--resources GPU=1"],
log_directory=f'./ce-perceptron/dask-logs',
job_script_prologue=["module load devel/cuda/12.1"],
job_extra_directives=["--gres=gpu:1"]
)
cluster.scale(5)
client = Client(cluster)
results = [client.submit(train, index, resources={"GPU": 1}) for index in indices]
My expected behaviour is that Dask submits five workers to the queue. Each worker takes a network to train with a given index, trains it on 16 cores and 1 GPU, and starts the next one when the training is finished. What happens, however, is that four workers are submitted to the queue, and only one of them starts to take networks and train them sequentially. The other workers are just idle.
I have tried increasing the number of processes, which I would think means running multiple network trainings on a single worker and splitting the resources. But this is also not correct as, in this case, it gives each process its own GPU despite the worker theoretically only having access to one. It also only runs on a single worker; the others are left idling.
I have also tried using map instead of submit. In this case, the workers die, or they try to run as many network trainings as possible on a single worker.
Finally, I have also tried using adapt, which is preferential to my workflow. However, when I do so, all of my workers keep dying with no logs produced in an endless cycle.
Even though I am reasonably familiar with clusters, especially SLURM clusters, as I mentioned above, I think I am missing something about how the API is supposed to work.
- 主要言語
- Python
- スター
- 256
- フォーク
- 149
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
dask/dask-jobqueue のほかの issue
-
bug LSF
難易度 3/5 1〜2日 初心者へのやさしさ 65/100
dask/dask-jobqueue#703 · コメント 1 件 ·
-
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dask/dask-jobqueue#701 ·
-
難易度 3/5 1〜2日 初心者へのやさしさ 45/100
dask/dask-jobqueue#699 · コメント 2 件 ·
-
bug
難易度 3/5 1〜2日 初心者へのやさしさ 38/100
dask/dask-jobqueue#692 · コメント 1 件 ·
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dask/dask-jobqueue#691 · コメント 7 件 ·
dask/dask-jobqueue の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
PedestrianDynamics/pyFDS-Evac#343 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
theskumar/python-dotenv#708 ·
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 2 日以内に返信
-
Docs Timedelta
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
pandas-dev/pandas#69919 ·
メンテナーはふだん 1 日以内に返信
-
API documentation
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
zephyrproject-rtos/west#1009 · コメント 2 件 ·
メンテナーはふだん 3 日以内に返信