Reusing intermediate results causes memory issues
@phofl がすでに取り組んでいます。
2024年2月14日 から。
評価
この issue はまだ評価されていません。
説明
Problem
Whenever we reuse intermediate results and there is a pipeline breaker (such as shuffles, joins, reductions, or groupby operations), it forces us to materialize the entire intermediate result (thus, it is breaking the pipelining that we could utilize for reuse for example between multiple element-wise operations).
This result materialization puts a hard limit on our ability to scale as I have observed in multiple TPC-H benchmark queries.
To illustrate this, run these two snippets on a cluster of your choice
With full intermediate result materialization
from dask_expr.datasets import timeseries
from distributed import Client
if __name__ == "__main__":
with Client() as client:
print(client.dashboard_link)
df = timeseries(start="2000-01-01", end="2020-12-31", freq="100ms", dtypes={"x": float})
# To compute mean, we have to fully materialize df, and we won't free its data
# until we have reused the chunks to compute a partial of the sum.
mean = df["x"].mean()
df[df["x"] > mean].sum().compute()
Without full intermediate result materialization
from dask_expr.datasets import timeseries
from distributed import Client
if __name__ == "__main__":
with Client() as client:
print(client.dashboard_link)
df = timeseries(start="2000-01-01", end="2020-12-31", freq="100ms", dtypes={"x": float})
# Compute the mean beforehand so that we don't have to keep all of `df` in memory
mean = df["x"].mean().compute()
df[df["x"] > mean].sum().compute()
Possible solution
The easiest approach would be to never reuse any intermediate results. This has a few downsides:
- Non-deterministic functions will lead to unexpected results
- We waste a lot of computational resources on recomputations
...but it will allow us to scale.
We can certainly get smarter about intermediate result materialization, but this will require some effort depending on how smart we want to be. (There's a body of (ongoing) research and implementations in the database world we could draw from.)
- 主要言語
- Python
- スター
- 89
- フォーク
- 26
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
dask/dask-expr のほかの issue
-
難易度 4/5 3〜5日 初心者へのやさしさ 38/100
-
難易度 3/5 1〜2日 初心者へのやさしさ 45/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 30/100
-
難易度 3/5 1〜2日 初心者へのやさしさ 25/100
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
kornia/kornia#5263 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
approved correction metadata
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
acl-org/acl-anthology#10133 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
BasedHardware/omi#20084 ·
メンテナーはふだん 1 日以内に返信
-
bug needs-acceptance wg/evaluation-quality
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
vllm-project/semantic-router#4424 ·
メンテナーはふだん 1 日以内に返信