Workers killed by signal 6/9 and timeout errors during cluster shutdown with `LocalCUDACluster`
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- distributed-systems
Hướng nghiên cứu
Payload không nêu tệp hay kiểm thử nào; hãy bắt đầu từ các entry point LocalCUDACluster và cluster.close và tái hiện quá trình tắt với cấu hình được cung cấp. Theo dõi tín hiệu và các thông báo timeout của distributed.nanny cùng với lỗi tcmalloc, sau đó xác định liệu đây có phải là hành vi được mong đợi hay cấu hình hoặc Best Practice nào mang lại quá trình tắt sạch sẽ.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
I'm using dask-cuda's LocalCUDACluster for GPU-based distributed computing in a Python script. While the computation completes successfully, I encounter multiple errors during the shutdown phase.
Specifically, after calling cluster.close() and attempting to gracefully shut down the Dask cluster, I see repeated logs like:
distributed.nanny - INFO - Worker process XXX was killed by signal 6
...
distributed.nanny - WARNING - Worker process still alive after 4.0 seconds, killing
...
distributed.nanny - INFO - Worker process XXX was killed by signal 9
Additionally, I get a traceback indicating a TimeoutError during internal cluster state correction:
tornado.application - ERROR - Exception in callback ...
TimeoutError
And finally, a memory-related error from tcmalloc:
src/tcmalloc.cc:284] Attempt to free invalid pointer 0x...
Environment Setup:
- Using
LocalCUDAClusterwith explicit GPU device configuration. - Disabled Dask optimizations (
optimization.fuse.active=False) and set conservative memory thresholds. - Workers are configured with
device_memory_limit="80GB"andthreads_per_worker=1. - Client and cluster are manually closed at the end of execution.
Code Snippet:
dask.config.set({"optimization.fuse.active": False})
dask.config.set({
"distributed.worker.memory.target": 0.6,
"distributed.worker.memory.spill": 0.7,
"distributed.worker.memory.pause": 0.8,
"distributed.worker.memory.terminate": 0.9,
"distributed.comm.timeouts.connect": "300s",
"distributed.comm.timeouts.tcp": "300s",
"distributed.worker.daemon": False,
"distributed.nanny.timeout": "60s"
})
cluster = LocalCUDACluster(
CUDA_VISIBLE_DEVICES=cuda_devices,
device_memory_limit="80GB",
n_workers=n_workers,
threads_per_worker=1,
dashboard_address=':0',
jit_unspill=False,
silence_logs=False
)
client = Client(cluster, timeout='60s')
client.wait_for_workers(n_workers, timeout=120)
# ... computation ...
cluster.close(timeout=300)
Expected Behavior:
Graceful shutdown of workers and scheduler without force-killing or timeout errors.
Actual Behavior:
Workers are terminated forcefully with signals 6 and 9, followed by timeout and memory-related errors during shutdown.
Environment:
- Dask version:
2024.12.1 - Dask-CUDA version:
25.2.0 - Python version: 3.12
- OS: Linux (assumed)
- Relevant packages:
cudf,cupy,torch,distributed, etc.
Question:
Is this expected behavior? Are there additional configurations or best practices to ensure clean shutdown of GPU clusters in Dask?
Any help or guidance would be greatly appreciated!
- Ngôn ngữ chính
- Python
- Star
- 1.7k
- Fork
- 778
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của dask/distributed
-
needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
dask/distributed#9366 ·
-
needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
dask/distributed#9353 ·
-
documentation
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 82/100
dask/distributed#8304 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
dask/distributed#4816 · 2 bình luận ·
-
documentation good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
dask/distributed#2378 · 2 bình luận ·
Tất cả issue của dask/distributed
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
qgis/QGIS-Plugins-Website#459 ·
-
bug severity:medium
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 2 ngày
-
bot-found bug priority: P3
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
madenvel/KalinkaPlayer#179 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
ls1intum/edutelligence#1098 ·
Maintainer thường phản hồi trong vòng 1 ngày