astronomy-commons/lsdb

Specify memory required to merge computed worker results

Open

#678 opened on Apr 4, 2025

 (1 comment) (0 reactions) (0 assignees)Python (21 forks)auto 404
enhancementhelp wanted

Repository metrics

Stars
 (54 stars)
PR merge metrics
 (PR metrics pending)

Description

When creating a Dask Client we specify the number of workers and the memory limit for each of them. This means that each worker is assigned the same amount of memory. E.g.:

dask.distributed.Client(n_workers=16, memory_limit="4GiB")

When we hit compute one of those workers is responsible for merging the results of all individual computations and, if the result is bigger than that single worker's memory, the computation fails. We should be able to specify a different amount of memory required for a final worker to complete the computation.

There seems to be no way of specifying which merge worker to use on compute of a Dask DataFrame but that seems to be possible through the Client: https://distributed.dask.org/en/stable/api.html#distributed.Client.persist. Requires some investigation.

Contributor guide