Reusing intermediate results causes memory issues
@phofl ya está trabajando en esto.
Desde el 14/2/2024.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Problem
Whenever we reuse intermediate results and there is a pipeline breaker (such as shuffles, joins, reductions, or groupby operations), it forces us to materialize the entire intermediate result (thus, it is breaking the pipelining that we could utilize for reuse for example between multiple element-wise operations).
This result materialization puts a hard limit on our ability to scale as I have observed in multiple TPC-H benchmark queries.
To illustrate this, run these two snippets on a cluster of your choice
With full intermediate result materialization
from dask_expr.datasets import timeseries
from distributed import Client
if __name__ == "__main__":
with Client() as client:
print(client.dashboard_link)
df = timeseries(start="2000-01-01", end="2020-12-31", freq="100ms", dtypes={"x": float})
# To compute mean, we have to fully materialize df, and we won't free its data
# until we have reused the chunks to compute a partial of the sum.
mean = df["x"].mean()
df[df["x"] > mean].sum().compute()
Without full intermediate result materialization
from dask_expr.datasets import timeseries
from distributed import Client
if __name__ == "__main__":
with Client() as client:
print(client.dashboard_link)
df = timeseries(start="2000-01-01", end="2020-12-31", freq="100ms", dtypes={"x": float})
# Compute the mean beforehand so that we don't have to keep all of `df` in memory
mean = df["x"].mean().compute()
df[df["x"] > mean].sum().compute()
Possible solution
The easiest approach would be to never reuse any intermediate results. This has a few downsides:
- Non-deterministic functions will lead to unexpected results
- We waste a lot of computational resources on recomputations
...but it will allow us to scale.
We can certainly get smarter about intermediate result materialization, but this will require some effort depending on how smart we want to be. (There's a body of (ongoing) research and implementations in the database world we could draw from.)
- Lenguaje dominante
- Python
- Estrellas
- 89
- Forks
- 26
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de dask/dask-expr
-
Dificultad 4/5 3-5 días Aptitud para principiantes 38/100
-
Dificultad 3/5 1-2 días Aptitud para principiantes 45/100
-
Predicate pull-up optimizationAbierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 30/100
-
Dificultad 4/5 3-5 días Aptitud para principiantes 30/100
-
Dificultad 3/5 1-2 días Aptitud para principiantes 25/100
Todos los issues de dask/dask-expr
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
kornia/kornia#5263 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Metadata correction for W16-5400Abiertoapproved correction metadata
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
acl-org/acl-anthology#10133 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
BasedHardware/omi#20084 ·
Los mantenedores suelen responder en 1 día
-
bug needs-acceptance wg/evaluation-quality
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
vllm-project/semantic-router#4424 ·
Los mantenedores suelen responder en 1 día