Keep exposure stores node-local in single-allocation runs; persist only per-exposure products
メンテナーはふだん 1 日以内に返信
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 18/100
- issue の種類
- リファクタリング
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- hpc, python
- 領域
- data, infrastructure
調査の方向性
Start with how sp run --in-allocation sets the exposure store location, and find where tile stores already use node-local $SLURM_TMPDIR. Then check sp_check_products.py and the star catalogue merge for any direct read of an exposure store. Done means exposure stores live on node-local disk in allocation mode, only the per-exposure PSF tars and exposure-map fragments are persisted, and a resumed failed batch still completes.
索引モデルが issue の本文から書いたものです。
説明
On Nibi a production batch is now one 144-core allocation running about 25 tiles (sp run --in-allocation). Its tile work is already node-local (#947 and the node-local tile stores). Its exposure stores are not: each batch builds about 43 of them on /scratch. Each node has 3 TB of local disk, so the exposure stores should live there too, and only the per-exposure products the catalogue needs should be persisted.
What the first production wave measured (smk-b01–b05, 10 Oct)
| batch | tiles | exposures | wall | billed core-h/tile |
|---|---|---|---|---|
| b01 | 25 | 42 | 2h09m | 12.4 |
| b02 | 25 | 43 | 1h58m | 11.3 |
| b03 | 25 | 42 | 2h08m | 12.3 |
| b04 | 25 | 43 | 2h06m | 12.1 |
| b05 | 25 | 44 | running |
- Scratch bytes are what binds. With wave 1's stores live,
/scratchreached ~1.9 TiB against Nibi's 1 TiB per-user soft quota, and it dropped back to 783 GiB as batches finished. The scale is ~6.6 GiB per exposure store, ~300 GiB per batch. Above the soft quota a 60-day grace counter runs, and it resets only when usage drops back below 1 TiB. So on/scratchwe can run only 2–3 batches at a time, or keep relying on the grace period. - File count does not bind at this scale. It is ~1,000 files per exposure store, ~50K per batch.
- Boundary recompute is real but cheap. The five batches build 214 exposure stores for 156 distinct exposures, so 27% are recomputed at batch edges. The exposure chain is ~1 of the ~12 billed core-h per tile (g12: 0.99), so the recompute costs about 2% of the total (inferred from g12's split).
Proposal
In --in-allocation mode, put the exposure stores on node-local disk ($SLURM_TMPDIR), next to the tile stores, and persist to the products directory only what downstream reads: the per-exposure PSF tars and the exposure-map fragments.
- A batch's
/scratchfootprint goes to roughly its products, so the soft quota stops limiting how many batches run at once. It is then limited by the scheduler and fairshare. - It makes #951 (packing per-exposure intermediates for the 1M-file quota) unnecessary for single-allocation runs. #951 still matters for the multi-job Slurm-executor launch.
- I/O moves off shared NFS. This is the same effect #947 had for
tile_detect: 8.7 → 1.0 min.
Not proposed: a cross-batch shared exposure store. It would save the ~2% recompute, but it would bring back a persistent shared store, with locking between concurrent batches and a bigger /scratch footprint. Ordering batches as compact 2D blocks gets most of that 2% for free.
What to check while implementing:
- Peak node-local use per batch. It should be ~300 GiB stores plus tile stores, well under 3 TB, but measure it on a dense batch.
- That a failed batch can be resumed. With node-local stores it re-runs from the start: ~2 h, which is acceptable.
- That nothing downstream (the star catalogue merge, the coverage and defect maps,
sp_check_products.py) reads an exposure store directly rather than a persisted product.
Claude (Opus, nibi chair) on behalf of Cail. Measurements by the nibi surveyor session from the smk-b01–b05 ledger.
- 主要言語
- Python
- スター
- 18
- フォーク
- 14
- 平均マージ
- 3日 22時間
- マージ済み PR(30日)
- 38
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
CosmoStat/shapepipe のほかの issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
メンテナーはふだん 1 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 18/100
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 52/100
CosmoStat/shapepipe#951 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
PSF star selection: survey other surveys' methods (MU_MAX−MAG_AUTO, colour)対応中かも @calumhrmurray が 4 日前に担当しました。 オープン
CosmoStat/shapepipe#950 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
CosmoStat/shapepipe の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
NousResearch/hermes-agent#136483 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last element対応中かも @peterdsharpe が今日担当しました。 オープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
pytorch/tensordict#2307 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
GrokModel.generate/a_generate pass an OpenAI-style list-of-dicts to xai_sdk.chat.user(), so every call crashes with a protobuf TypeError before any network I/O対応中かも @Christian-Sidak が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
confident-ai/deepeval#3436 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信