Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Workers killed by signal 6/9 and timeout errors during cluster shutdown with `LocalCUDACluster`

オープン
#9,100 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
python

調査の方向性

ペイロードにはファイルもテストも記載されていないため、LocalCUDACluster と cluster.close のエントリーポイントから始め、提供された設定でシャットダウンを再現してください。distributed.nanny のシグナルおよびタイムアウトメッセージを tcmalloc エラーと併せて追跡し、その後、これが想定された動作なのか、またはどの設定や Best Practice によって正常なシャットダウンが実現するのかを明らかにしてください。

索引モデルが issue の本文から書いたものです。

説明

needs triage

I'm using dask-cuda's LocalCUDACluster for GPU-based distributed computing in a Python script. While the computation completes successfully, I encounter multiple errors during the shutdown phase.

Specifically, after calling cluster.close() and attempting to gracefully shut down the Dask cluster, I see repeated logs like:

distributed.nanny - INFO - Worker process XXX was killed by signal 6
...
distributed.nanny - WARNING - Worker process still alive after 4.0 seconds, killing
...
distributed.nanny - INFO - Worker process XXX was killed by signal 9

Additionally, I get a traceback indicating a TimeoutError during internal cluster state correction:

tornado.application - ERROR - Exception in callback ...
TimeoutError

And finally, a memory-related error from tcmalloc:

src/tcmalloc.cc:284] Attempt to free invalid pointer 0x...

Environment Setup:

  • Using LocalCUDACluster with explicit GPU device configuration.
  • Disabled Dask optimizations (optimization.fuse.active=False) and set conservative memory thresholds.
  • Workers are configured with device_memory_limit="80GB" and threads_per_worker=1.
  • Client and cluster are manually closed at the end of execution.

Code Snippet:

dask.config.set({"optimization.fuse.active": False})
dask.config.set({
    "distributed.worker.memory.target": 0.6,
    "distributed.worker.memory.spill": 0.7,
    "distributed.worker.memory.pause": 0.8,
    "distributed.worker.memory.terminate": 0.9,
    "distributed.comm.timeouts.connect": "300s",
    "distributed.comm.timeouts.tcp": "300s",
    "distributed.worker.daemon": False,
    "distributed.nanny.timeout": "60s"
})

cluster = LocalCUDACluster(
    CUDA_VISIBLE_DEVICES=cuda_devices,
    device_memory_limit="80GB",
    n_workers=n_workers,
    threads_per_worker=1,
    dashboard_address=':0',
    jit_unspill=False,
    silence_logs=False
)

client = Client(cluster, timeout='60s')
client.wait_for_workers(n_workers, timeout=120)

# ... computation ...

cluster.close(timeout=300)

Expected Behavior:

Graceful shutdown of workers and scheduler without force-killing or timeout errors.

Actual Behavior:

Workers are terminated forcefully with signals 6 and 9, followed by timeout and memory-related errors during shutdown.

Environment:

  • Dask version: 2024.12.1
  • Dask-CUDA version: 25.2.0
  • Python version: 3.12
  • OS: Linux (assumed)
  • Relevant packages: cudf, cupy, torch, distributed, etc.

Question:

Is this expected behavior? Are there additional configurations or best practices to ensure clean shutdown of GPU clusters in Dask?

Any help or guidance would be greatly appreciated!

主要言語
Python
スター
1.7k
フォーク
778
PR マージ指標
30日以内にマージされた PR はありません

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

dask/distributed のほかの issue

dask/distributed の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。