Run, observe, and analyze Slurm benchmarks
メンテナーはふだん 1 日以内に返信
@andreatnvidia がすでに取り組んでいます。
2026年8月24日 から。
評価
この issue はまだ評価されていません。
説明
Priority Level
High
Task Summary
Implement deterministic benchmark expansion, batch execution, fresh-process observation, and analysis for the optional Slurm integration.
Technical Details & Implementation Plan
- Expand strict benchmark intent into deterministic child run configurations and immutable benchmark records.
- Submit child runs through the public Slurm execution service and return without resident monitoring.
- Persist benchmark-to-run identity so later processes can observe and analyze the same children.
- Refresh child state through normalized scheduler and persisted-state evidence.
- Compute stable aggregate analysis from validated child results while preserving failed, incomplete, and missing-run classifications.
- Expose equivalent Python service and CLI operations for benchmark run and analyze workflows.
Acceptance criteria
- Equivalent benchmark input expands to the same ordered child runs and digests.
- Benchmark submission returns after scheduling children and does not require a resident controller.
- Analysis works from a fresh process and never guesses success from incomplete evidence.
- Missing, failed, partial, stale, and scheduler-inconsistent child runs remain explicit in results.
- Local/fake tests cover expansion, submission, refresh, mixed outcomes, and deterministic analysis.
- Real-cluster acceptance validates at least one multi-run benchmark workflow before release.
Out of scope
- Interactive dashboards or resident monitoring.
- New benchmark algorithms unrelated to Slurm execution.
- Generic scheduler or platform adapters.
Investigation / Context
This is the benchmark implementation lane in #850. #865 and #872 explicitly exclude benchmark implementation, while #870 treats benchmark analysis as a final acceptance scenario.
Agent Plan / Findings
Reuse the same immutable planning, execution, and state contracts as ordinary runs; benchmark records should add hierarchy and analysis intent rather than a second control plane.
Dependencies
Depends on shared benchmark records in #873, fake infrastructure in #872, deterministic planning in #875, the client worker in #876, the public service foundation and operational run/observe capabilities from 874#1 and 874#2, the one-node runtime capability from 868#2, and the persistence, winner, observation, and reconciliation capabilities from 869#1, 869#2, and 869#3. It does not depend on distributed/failure hardening in 868#3, retry/collection in 869#4, or the later 874#3, 874#4, and 874#5 slices.
- 主要言語
- Python
- スター
- 2.3k
- フォーク
- 211
- 平均マージ
- 3日 20時間
- マージ済み PR(30日)
- 40
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA-NeMo/DataDesigner のほかの issue
-
task
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA-NeMo/DataDesigner#760 ·
メンテナーはふだん 1 日以内に返信
-
task
難易度 4/5 3〜5日 初心者へのやさしさ 42/100
NVIDIA-NeMo/DataDesigner#964 ·
メンテナーはふだん 1 日以内に返信
-
enhancement triaged
難易度 4/5 3〜5日 初心者へのやさしさ 40/100
NVIDIA-NeMo/DataDesigner#956 ·
メンテナーはふだん 1 日以内に返信
-
task
難易度 5/5 1週間以上 初心者へのやさしさ 42/100
NVIDIA-NeMo/DataDesigner#947 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 52/100
NVIDIA-NeMo/DataDesigner#946 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
NVIDIA-NeMo/DataDesigner の issue をすべて見る
似ている issue
-
ACK_WAITING HELP_WANTED UPDATE_CS
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
OWASP/CheatSheetSeries#2458 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
BasedHardware/omi#19711 ·
メンテナーはふだん 1 日以内に返信
-
Qwen3_5MoeModel no longer returns router_logits, breaking aux loss with output_router_logits=Trueオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
huggingface/transformers#49172 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
vllm-project/vllm-metal#885 ·
メンテナーはふだん 1 日以内に返信