astronomy-commons/lsdb

Specify memory required to merge computed worker results

Ouverte

#678 ouverte le 4 avr. 2025

 (1 commentaire) (0 réaction) (0 personne assignée)Python (21 forks)auto 404
enhancementhelp wanted

Métriques du dépôt

Stars
 (54 étoiles)
Métriques de merge PR
 (Métriques PR en attente)

Description

When creating a Dask Client we specify the number of workers and the memory limit for each of them. This means that each worker is assigned the same amount of memory. E.g.:

dask.distributed.Client(n_workers=16, memory_limit="4GiB")

When we hit compute one of those workers is responsible for merging the results of all individual computations and, if the result is bigger than that single worker's memory, the computation fails. We should be able to specify a different amount of memory required for a final worker to complete the computation.

There seems to be no way of specifying which merge worker to use on compute of a Dask DataFrame but that seems to be possible through the Client: https://distributed.dask.org/en/stable/api.html#distributed.Client.persist. Requires some investigation.

Guide contributeur