Reusing intermediate results causes memory issues
@phofl ci sta già lavorando.
Dal 14/2/2024.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Problem
Whenever we reuse intermediate results and there is a pipeline breaker (such as shuffles, joins, reductions, or groupby operations), it forces us to materialize the entire intermediate result (thus, it is breaking the pipelining that we could utilize for reuse for example between multiple element-wise operations).
This result materialization puts a hard limit on our ability to scale as I have observed in multiple TPC-H benchmark queries.
To illustrate this, run these two snippets on a cluster of your choice
With full intermediate result materialization
from dask_expr.datasets import timeseries
from distributed import Client
if __name__ == "__main__":
with Client() as client:
print(client.dashboard_link)
df = timeseries(start="2000-01-01", end="2020-12-31", freq="100ms", dtypes={"x": float})
# To compute mean, we have to fully materialize df, and we won't free its data
# until we have reused the chunks to compute a partial of the sum.
mean = df["x"].mean()
df[df["x"] > mean].sum().compute()
Without full intermediate result materialization
from dask_expr.datasets import timeseries
from distributed import Client
if __name__ == "__main__":
with Client() as client:
print(client.dashboard_link)
df = timeseries(start="2000-01-01", end="2020-12-31", freq="100ms", dtypes={"x": float})
# Compute the mean beforehand so that we don't have to keep all of `df` in memory
mean = df["x"].mean().compute()
df[df["x"] > mean].sum().compute()
Possible solution
The easiest approach would be to never reuse any intermediate results. This has a few downsides:
- Non-deterministic functions will lead to unexpected results
- We waste a lot of computational resources on recomputations
...but it will allow us to scale.
We can certainly get smarter about intermediate result materialization, but this will require some effort depending on how smart we want to be. (There's a body of (ongoing) research and implementations in the database world we could draw from.)
- Lingua principale
- Python
- Stelle
- 89
- Fork
- 26
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di dask/dask-expr
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 38/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 45/100
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 30/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 25/100
Tutte le issue di dask/dask-expr
Issue simili
-
repo-audit
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
scverse/repo-health#20 ·
I maintainer di solito rispondono entro 1 giorno
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
phasespace-labs/palinode#232 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
collective/icalendar#1858 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
lfx-mcp cannot supply global variables: LangflowClient drops X-LANGFLOW-GLOBAL-VAR-* from envApertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
langflow-ai/langflow#15496 ·
I maintainer di solito rispondono entro 1 giorno