Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Run, observe, and analyze Slurm benchmarks

Đã đóng
#877 1 bình luận 0 reaction 1 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

@andreatnvidia đang làm issue này rồi.

Từ ngày 24/8/2026.

Đánh giá

Issue này chưa được đánh giá.

Mô tả

task
Priority Level

High

Task Summary

Implement deterministic benchmark expansion, batch execution, fresh-process observation, and analysis for the optional Slurm integration.

Technical Details & Implementation Plan
  • Expand strict benchmark intent into deterministic child run configurations and immutable benchmark records.
  • Submit child runs through the public Slurm execution service and return without resident monitoring.
  • Persist benchmark-to-run identity so later processes can observe and analyze the same children.
  • Refresh child state through normalized scheduler and persisted-state evidence.
  • Compute stable aggregate analysis from validated child results while preserving failed, incomplete, and missing-run classifications.
  • Expose equivalent Python service and CLI operations for benchmark run and analyze workflows.
Acceptance criteria
  • Equivalent benchmark input expands to the same ordered child runs and digests.
  • Benchmark submission returns after scheduling children and does not require a resident controller.
  • Analysis works from a fresh process and never guesses success from incomplete evidence.
  • Missing, failed, partial, stale, and scheduler-inconsistent child runs remain explicit in results.
  • Local/fake tests cover expansion, submission, refresh, mixed outcomes, and deterministic analysis.
  • Real-cluster acceptance validates at least one multi-run benchmark workflow before release.
Out of scope
  • Interactive dashboards or resident monitoring.
  • New benchmark algorithms unrelated to Slurm execution.
  • Generic scheduler or platform adapters.
Investigation / Context

This is the benchmark implementation lane in #850. #865 and #872 explicitly exclude benchmark implementation, while #870 treats benchmark analysis as a final acceptance scenario.

Agent Plan / Findings

Reuse the same immutable planning, execution, and state contracts as ordinary runs; benchmark records should add hierarchy and analysis intent rather than a second control plane.

Dependencies

Depends on shared benchmark records in #873, fake infrastructure in #872, deterministic planning in #875, the client worker in #876, the public service foundation and operational run/observe capabilities from 874#1 and 874#2, the one-node runtime capability from 868#2, and the persistence, winner, observation, and reconciliation capabilities from 869#1, 869#2, and 869#3. It does not depend on distributed/failure hardening in 868#3, retry/collection in 869#4, or the later 874#3, 874#4, and 874#5 slices.

Ngôn ngữ chính
Python
Star
2.3k
Fork
211
Merge trung bình
3 ngày 12 giờ
Pull request đã merge (30 ngày)
45

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của NVIDIA-NeMo/DataDesigner

Tất cả issue của NVIDIA-NeMo/DataDesigner

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.