done_callback raises RuntimeError during interpreter shutdown
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 50/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- distributed-systems
Hướng nghiên cứu
Bắt đầu trong client.py tại done_callback và lần theo đường đi callback(future) đến _cb_executor.submit(execute_callback). Chạy ví dụ tối thiểu hoàn chỉnh có thể kiểm chứng trong khi trình thông dịch đang tắt; done có nghĩa là RuntimeError liên quan đến việc tắt được bỏ qua, trong khi các trường hợp RuntimeError khác vẫn xuất hiện.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the issue:
done_callback in client.py raises RuntimeError: cannot schedule new futures after interpreter shutdown when futures complete during interpreter shutdown.
Future.add_done_callback() schedules callbacks via Future._cb_executor (a class-level ThreadPoolExecutor). During shutdown, CPython's concurrent.futures.thread._python_exit() atexit handler kills all ThreadPoolExecutor instances and sets a module-level _shutdown = True flag. The Tornado IOLoop daemon thread is still alive draining done_callback coroutines — when they call _cb_executor.submit(), it raises RuntimeError.
The fix is to catch shutdown-related RuntimeError in done_callback since the callback is no longer meaningful if the interpreter is shutting down.
Minimal Complete Verifiable Example:
import time
from distributed import Client, LocalCluster
def on_done(future):
print(f"Callback: {future.key}")
if __name__ == "__main__":
cluster = LocalCluster(n_workers=2, threads_per_worker=1)
client = Client(cluster)
# pure=False ensures unique keys; staggered sleeps so futures complete during shutdown
futures = client.map(time.sleep, [0.2 * i for i in range(20)], pure=False)
for f in futures:
f.add_done_callback(on_done)
# Exit while futures are still completing — don't call client.close()
time.sleep(1.5)
Anything else we need to know?:
The error path is: done_callback → callback(future) → _cb_executor.submit(execute_callback) → RuntimeError.
Even creating a new ThreadPoolExecutor would not help because CPython's concurrent.futures.thread.submit() checks both the instance flag (self._shutdown) and the module-level flag (_shutdown), which was already set by _python_exit().
Suggested fix in done_callback:
async def done_callback(future, callback):
while future.status == "pending":
await future._state.wait()
try:
callback(future)
except RuntimeError as e:
if "shutdown" not in str(e) and "interpreter" not in str(e):
raise
This is also a prerequisite for a related dask-gateway bug where cleanup_lingering_clusters atexit fails to send HTTP DELETE to shut down clusters because the same executor shutdown blocks aiohttp's DNS resolution.
Environment:
- Dask version: 2025.9.1
- Python version: 3.11.14
- Operating System: macOS 15.7.4 (arm64)
- Install method: uv (pip)
- Ngôn ngữ chính
- Python
- Star
- 1.7k
- Fork
- 778
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của dask/distributed
-
needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
dask/distributed#9366 ·
-
needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
dask/distributed#9353 ·
-
documentation
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 82/100
dask/distributed#8304 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
dask/distributed#4816 · 2 bình luận ·
-
documentation good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
dask/distributed#2378 · 2 bình luận ·
Tất cả issue của dask/distributed
Issue tương tự
-
bug status/needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
prowler-cloud/prowler#12887 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
use-agent-os/agent-os#3484 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
open-telemetry/opentelemetry-python-contrib#5113 · 2 bình luận · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
external
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
langchain-ai/docs#6255 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày