Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Workers killed by signal 6/9 and timeout errors during cluster shutdown with `LocalCUDACluster`

Đang mở
#9,100 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
35/100
Loại issue
Lỗi
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
python
Lĩnh vực
distributed-systems

Hướng nghiên cứu

Payload không nêu tệp hay kiểm thử nào; hãy bắt đầu từ các entry point LocalCUDACluster và cluster.close và tái hiện quá trình tắt với cấu hình được cung cấp. Theo dõi tín hiệu và các thông báo timeout của distributed.nanny cùng với lỗi tcmalloc, sau đó xác định liệu đây có phải là hành vi được mong đợi hay cấu hình hoặc Best Practice nào mang lại quá trình tắt sạch sẽ.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

needs triage

I'm using dask-cuda's LocalCUDACluster for GPU-based distributed computing in a Python script. While the computation completes successfully, I encounter multiple errors during the shutdown phase.

Specifically, after calling cluster.close() and attempting to gracefully shut down the Dask cluster, I see repeated logs like:

distributed.nanny - INFO - Worker process XXX was killed by signal 6
...
distributed.nanny - WARNING - Worker process still alive after 4.0 seconds, killing
...
distributed.nanny - INFO - Worker process XXX was killed by signal 9

Additionally, I get a traceback indicating a TimeoutError during internal cluster state correction:

tornado.application - ERROR - Exception in callback ...
TimeoutError

And finally, a memory-related error from tcmalloc:

src/tcmalloc.cc:284] Attempt to free invalid pointer 0x...

Environment Setup:

  • Using LocalCUDACluster with explicit GPU device configuration.
  • Disabled Dask optimizations (optimization.fuse.active=False) and set conservative memory thresholds.
  • Workers are configured with device_memory_limit="80GB" and threads_per_worker=1.
  • Client and cluster are manually closed at the end of execution.

Code Snippet:

dask.config.set({"optimization.fuse.active": False})
dask.config.set({
    "distributed.worker.memory.target": 0.6,
    "distributed.worker.memory.spill": 0.7,
    "distributed.worker.memory.pause": 0.8,
    "distributed.worker.memory.terminate": 0.9,
    "distributed.comm.timeouts.connect": "300s",
    "distributed.comm.timeouts.tcp": "300s",
    "distributed.worker.daemon": False,
    "distributed.nanny.timeout": "60s"
})

cluster = LocalCUDACluster(
    CUDA_VISIBLE_DEVICES=cuda_devices,
    device_memory_limit="80GB",
    n_workers=n_workers,
    threads_per_worker=1,
    dashboard_address=':0',
    jit_unspill=False,
    silence_logs=False
)

client = Client(cluster, timeout='60s')
client.wait_for_workers(n_workers, timeout=120)

# ... computation ...

cluster.close(timeout=300)

Expected Behavior:

Graceful shutdown of workers and scheduler without force-killing or timeout errors.

Actual Behavior:

Workers are terminated forcefully with signals 6 and 9, followed by timeout and memory-related errors during shutdown.

Environment:

  • Dask version: 2024.12.1
  • Dask-CUDA version: 25.2.0
  • Python version: 3.12
  • OS: Linux (assumed)
  • Relevant packages: cudf, cupy, torch, distributed, etc.

Question:

Is this expected behavior? Are there additional configurations or best practices to ensure clean shutdown of GPU clusters in Dask?

Any help or guidance would be greatly appreciated!

Ngôn ngữ chính
Python
Star
1.7k
Fork
778
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của dask/distributed

Tất cả issue của dask/distributed

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.