Consider using object pools for internal state keeping
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 20/100
- Issue 类型
- 重构
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
- 技术栈
- python
- 领域
- performance
调研方向
该 issue 未列出源文件、入口点或测试;请先对 scheduler 中的 set 分配进行性能分析,并测量其对内存和运行时的影响。将这些测量结果与提议的 SetPool/PooledSet 概念进行比较,只有在项目测试中证明存在可测量的收益和安全的复用行为后,才将这项工作视为完成。
由索引模型根据 Issue 内容生成。
描述
Dask is creating many small objects and not just for very large graphs but this is pretty much a built in thing. This is one of not the most prominent reason why dask was originally built as a "tuple of tuple of tuple ..." machinery. Even the scheduler internally only adopted the usage of custom classes only a couple of years back because instantiation can be very costly. However, modern python versions have gotten much better at managing overhead so this is negligible for most classes.
The one type of objects we're still affected by, both in terms of memory but also in terms of runtime are sets. Yes, sets! (I might share profiles but this is mostly an issue to preseve an idea). One way to work around instantiation cost of sets is to use an object pool design. Effectively, we'd "disable" garbage collection and would resurrect objects on finalization in a way that would allow us to reuse them.
A minimal version of this would look like that
import sys
class SetPool:
def __init__(self):
self._sets = []
def add(self, obj):
obj.clear()
self._sets.append(obj)
def get(self):
try:
new = self._sets.pop()
return new
except:
return None
def stored_size(self):
return sum(map(sys.getsizeof, self._sets))
globalpool = SetPool()
class PooledSet(set):
def __new__(cls, *iterables):
obj = globalpool.get()
if obj is not None:
return obj
return super().__new__(cls, *iterables)
def __del__(self):
print(f"Resurrecting {id(self)}")
globalpool.add(self)
This could be expanded on need, e.g. by hinting towards whether this would be an empty set, a very large set or a somewhat normal one.
Haven't tried out what the actual impact would be but it is a fun concept that could help if we actually want/need to optimize for this
- 主要语言
- Python
- 星标
- 1.7k
- 派生
- 780
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
dask/distributed 的其他 Issue
-
needs triage
难度 2/5 1-3 小时 新手友好度 72/100
dask/distributed#9366 ·
-
needs triage
难度 2/5 1-3 小时 新手友好度 84/100
dask/distributed#9353 ·
-
documentation
难度 1/5 1-3 小时 新手友好度 82/100
dask/distributed#8304 ·
-
难度 2/5 1-3 小时 新手友好度 74/100
dask/distributed#4816 · 2 条评论 ·
-
documentation good first issue
难度 2/5 1-3 小时 新手友好度 74/100
dask/distributed#2378 · 2 条评论 ·
相似的 Issue
-
repo-audit
难度 2/5 1-3 小时 新手友好度 75/100
scverse/repo-health#20 ·
维护者通常 1 天内回复
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memories未关闭
难度 2/5 1-3 小时 新手友好度 85/100
phasespace-labs/palinode#232 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
collective/icalendar#1858 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
langflow-ai/langflow#15496 ·
维护者通常 1 天内回复