Workers killed by signal 6/9 and timeout errors during cluster shutdown with `LocalCUDACluster`
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 停滞
- 技術スタック
- python
調査の方向性
ペイロードにはファイルもテストも記載されていないため、LocalCUDACluster と cluster.close のエントリーポイントから始め、提供された設定でシャットダウンを再現してください。distributed.nanny のシグナルおよびタイムアウトメッセージを tcmalloc エラーと併せて追跡し、その後、これが想定された動作なのか、またはどの設定や Best Practice によって正常なシャットダウンが実現するのかを明らかにしてください。
索引モデルが issue の本文から書いたものです。
説明
I'm using dask-cuda's LocalCUDACluster for GPU-based distributed computing in a Python script. While the computation completes successfully, I encounter multiple errors during the shutdown phase.
Specifically, after calling cluster.close() and attempting to gracefully shut down the Dask cluster, I see repeated logs like:
distributed.nanny - INFO - Worker process XXX was killed by signal 6
...
distributed.nanny - WARNING - Worker process still alive after 4.0 seconds, killing
...
distributed.nanny - INFO - Worker process XXX was killed by signal 9
Additionally, I get a traceback indicating a TimeoutError during internal cluster state correction:
tornado.application - ERROR - Exception in callback ...
TimeoutError
And finally, a memory-related error from tcmalloc:
src/tcmalloc.cc:284] Attempt to free invalid pointer 0x...
Environment Setup:
- Using
LocalCUDAClusterwith explicit GPU device configuration. - Disabled Dask optimizations (
optimization.fuse.active=False) and set conservative memory thresholds. - Workers are configured with
device_memory_limit="80GB"andthreads_per_worker=1. - Client and cluster are manually closed at the end of execution.
Code Snippet:
dask.config.set({"optimization.fuse.active": False})
dask.config.set({
"distributed.worker.memory.target": 0.6,
"distributed.worker.memory.spill": 0.7,
"distributed.worker.memory.pause": 0.8,
"distributed.worker.memory.terminate": 0.9,
"distributed.comm.timeouts.connect": "300s",
"distributed.comm.timeouts.tcp": "300s",
"distributed.worker.daemon": False,
"distributed.nanny.timeout": "60s"
})
cluster = LocalCUDACluster(
CUDA_VISIBLE_DEVICES=cuda_devices,
device_memory_limit="80GB",
n_workers=n_workers,
threads_per_worker=1,
dashboard_address=':0',
jit_unspill=False,
silence_logs=False
)
client = Client(cluster, timeout='60s')
client.wait_for_workers(n_workers, timeout=120)
# ... computation ...
cluster.close(timeout=300)
Expected Behavior:
Graceful shutdown of workers and scheduler without force-killing or timeout errors.
Actual Behavior:
Workers are terminated forcefully with signals 6 and 9, followed by timeout and memory-related errors during shutdown.
Environment:
- Dask version:
2024.12.1 - Dask-CUDA version:
25.2.0 - Python version: 3.12
- OS: Linux (assumed)
- Relevant packages:
cudf,cupy,torch,distributed, etc.
Question:
Is this expected behavior? Are there additional configurations or best practices to ensure clean shutdown of GPU clusters in Dask?
Any help or guidance would be greatly appreciated!
- 主要言語
- Python
- スター
- 1.7k
- フォーク
- 778
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
dask/distributed のほかの issue
-
needs triage
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
dask/distributed#9366 ·
-
needs triage
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
dask/distributed#9353 ·
-
documentation
難易度 1/5 1〜3時間 初心者へのやさしさ 82/100
dask/distributed#8304 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
dask/distributed#4816 · コメント 2 件 ·
-
documentation good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
dask/distributed#2378 · コメント 2 件 ·
dask/distributed の issue をすべて見る
似ている issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
メンテナーはふだん 1 日以内に返信
-
instance instance add
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
searxng/searx-instances#941 · コメント 1 件 ·
-
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
FluidNumerics/fluid-walk-blocker#89 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信