astronomy-commons/lsdb

Specify memory required to merge computed worker results

オープン

#678 opened on 2025/04/04

 (1 件のコメント) (0 件のリアクション) (0 人の担当者)Python (21 件のフォーク)auto 404
enhancementhelp wanted

Repository metrics

Stars
 (54 個のスター)
PR merge metrics
 (PR metrics pending)

説明

When creating a Dask Client we specify the number of workers and the memory limit for each of them. This means that each worker is assigned the same amount of memory. E.g.:

dask.distributed.Client(n_workers=16, memory_limit="4GiB")

When we hit compute one of those workers is responsible for merging the results of all individual computations and, if the result is bigger than that single worker's memory, the computation fails. We should be able to specify a different amount of memory required for a final worker to complete the computation.

There seems to be no way of specifying which merge worker to use on compute of a Dask DataFrame but that seems to be possible through the Client: https://distributed.dask.org/en/stable/api.html#distributed.Client.persist. Requires some investigation.

コントリビューターガイド