SpecCluster._correct_state_internal closes workers with Nanny's default 5s timeout; shutdown often fails on LocalCluster(processes=True)
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 48/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- python
调研方向
traceback 指向 distributed/deploy/spec.py 和 SpecCluster._correct_state_internal();首先,在保留已完成 futures 的同时,使用 LocalCluster(processes=True) 和 client.shutdown() 示例复现该问题。跟踪 retirement 和 AMM 路径,然后建立回归检查,证明完整集群 teardown 可以避免不必要的 retirement 工作,并且不再引发 TimeoutError。
由索引模型根据 Issue 内容生成。
描述
TimeoutError during client.shutdown() due to unnecessary AMM retirement when futures are not explicitly released
Description
When shutting down a dedicated LocalCluster via client.shutdown() after batch processing is complete, Dask still executes the full graceful worker retirement path (retire_workers + Active Memory Manager replication/drop), even though all workers are being removed and the results are no longer needed.
If the client has not released its futures first, completed task data remains in the scheduler's state memory on the workers. During shutdown, the AMM attempts to replicate or drop those keys across workers that are simultaneously shutting down. This causes unnecessary network/memory overhead and slows down worker teardown. Consequently, this can exceed the Nanny’s default process.join timeout (~4s) and surface as a TimeoutError / Tornado ERROR log during SpecCluster._correct_state_internal().
Releasing all client-held keys before cluster.close() avoids the problem in practice, which suggests the current shutdown path is doing redundant work that a full-cluster teardown should not require.
Sample Code from My Project
Create cluster and client
cls.cluster = LocalCluster(
name="cluster",
n_workers=cls.n_workers,
threads_per_worker=cls.threads_per_worker,
memory_limit=0,
dashboard_address=dashboard_address
)
cls.client = Client(
name="client",
address=cls.cluster,
direct_to_workers=True
)
Shutdown
@classmethod
def system_teardown(cls):
with suppress(Exception):
cls.client.shutdown()
Error Stack Trace
Occasionally, the following error occurs during shutdown:
2026-06-10 14:56:47,407 - tornado.application - ERROR - Exception in callback functools.partial(<bound method IOLoop._discard_future_result of <tornado.platform.asyncio.AsyncIOMainLoop object at 0x00000288583A1D30>>, <Task finished name='Task-1174981' coro=<Spec Cluster._correct_state_internal() done, defined at .venvLibsite-packagesdistributeddeployspec.py:346> exception=TimeoutError()>)
Traceback (most recent call last):
File ".venvLibsite-packagesdistributedutils.py", line 1910, in wait_for
return await fut
asyncio.exceptions.CancelledError
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File ".venvLibsite-packagestornadoioloop.py", line 758, in _run_callback
ret = callback()
File ".venvLibsite-packagestornadoioloop.py", line 782, in _discard_future_result
future.result()
TimeoutError
Root Cause Analysis
When client.shutdown() is called without prior key release, completed task data remains in the scheduler's state memory on the workers. During the shutdown sequence, the Active Memory Manager (AMM) attempts to replicate or drop these keys across workers that are concurrently being torn down. This triggers unnecessary network/memory operations and significantly slows down the worker teardown process. Ultimately, this exceeds the Nanny’s default process.join timeout (~4s), resulting in a TimeoutError within SpecCluster._correct_state_internal().
Suggestions / Recommendations
Optimize Full-Cluster Teardown: Consider bypassing or short-circuiting the AMM retirement logic (retire_workers, data replication/drop) when a full-cluster teardown is initiated via client.shutdown(). Since all workers are being removed anyway, preserving or migrating their data is redundant.
Auto-Release on Shutdown: Alternatively, automatically release all client-held keys before initiating the shutdown sequence to prevent the AMM from acting on stale future references.
Documentation: In the meantime, it might be helpful to document this behavior and recommend users explicitly release their futures (e.g., via client.release() or iterating over client.futures) before calling shutdown() as a best practice.
Thanks!
- 主要语言
- Python
- 星标
- 1.7k
- 派生
- 780
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
dask/distributed 的其他 Issue
-
needs triage
难度 2/5 1-3 小时 新手友好度 72/100
dask/distributed#9366 ·
-
needs triage
难度 2/5 1-3 小时 新手友好度 84/100
dask/distributed#9353 ·
-
documentation
难度 1/5 1-3 小时 新手友好度 82/100
dask/distributed#8304 ·
-
难度 2/5 1-3 小时 新手友好度 74/100
dask/distributed#4816 · 2 条评论 ·
-
documentation good first issue
难度 2/5 1-3 小时 新手友好度 74/100
dask/distributed#2378 · 2 条评论 ·
相似的 Issue
-
adr
难度 2/5 1-3 小时 新手友好度 72/100
kristofdegrave/homeassistant-smart-charging#1607 ·
维护者通常 1 天内回复
-
namespace operations
难度 2/5 1-3 小时 新手友好度 64/100
EclipseFdn/open-vsx.org#13665 ·
维护者通常 1 天内回复
-
doc good first issue help wanted
难度 2/5 1-3 小时 新手友好度 68/100
collective/icalendar#1865 · 2 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
canonical/opentelemetry-collector-operator#409 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 85/100
mozilla/addons-release-tests#1243 ·
维护者通常 1 天内回复