Run, observe, and analyze Slurm benchmarks
Los mantenedores suelen responder en 1 día
@andreatnvidia ya está trabajando en esto.
Desde el 24/8/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Priority Level
High
Task Summary
Implement deterministic benchmark expansion, batch execution, fresh-process observation, and analysis for the optional Slurm integration.
Technical Details & Implementation Plan
- Expand strict benchmark intent into deterministic child run configurations and immutable benchmark records.
- Submit child runs through the public Slurm execution service and return without resident monitoring.
- Persist benchmark-to-run identity so later processes can observe and analyze the same children.
- Refresh child state through normalized scheduler and persisted-state evidence.
- Compute stable aggregate analysis from validated child results while preserving failed, incomplete, and missing-run classifications.
- Expose equivalent Python service and CLI operations for benchmark run and analyze workflows.
Acceptance criteria
- Equivalent benchmark input expands to the same ordered child runs and digests.
- Benchmark submission returns after scheduling children and does not require a resident controller.
- Analysis works from a fresh process and never guesses success from incomplete evidence.
- Missing, failed, partial, stale, and scheduler-inconsistent child runs remain explicit in results.
- Local/fake tests cover expansion, submission, refresh, mixed outcomes, and deterministic analysis.
- Real-cluster acceptance validates at least one multi-run benchmark workflow before release.
Out of scope
- Interactive dashboards or resident monitoring.
- New benchmark algorithms unrelated to Slurm execution.
- Generic scheduler or platform adapters.
Investigation / Context
This is the benchmark implementation lane in #850. #865 and #872 explicitly exclude benchmark implementation, while #870 treats benchmark analysis as a final acceptance scenario.
Agent Plan / Findings
Reuse the same immutable planning, execution, and state contracts as ordinary runs; benchmark records should add hierarchy and analysis intent rather than a second control plane.
Dependencies
Depends on shared benchmark records in #873, fake infrastructure in #872, deterministic planning in #875, the client worker in #876, the public service foundation and operational run/observe capabilities from 874#1 and 874#2, the one-node runtime capability from 868#2, and the persistence, winner, observation, and reconciliation capabilities from 869#1, 869#2, and 869#3. It does not depend on distributed/failure hardening in 868#3, retry/collection in 869#4, or the later 874#3, 874#4, and 874#5 slices.
- Lenguaje dominante
- Python
- Estrellas
- 2.3k
- Forks
- 211
- Merge medio
- 3 d 20 h
- PR fusionados (30 d)
- 40
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA-NeMo/DataDesigner
-
task
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Los mantenedores suelen responder en 1 día
-
enhancement triaged
Dificultad 4/5 3-5 días Aptitud para principiantes 40/100
NVIDIA-NeMo/DataDesigner#956 ·
Los mantenedores suelen responder en 1 día
-
task
Dificultad 5/5 Más de una semana Aptitud para principiantes 42/100
NVIDIA-NeMo/DataDesigner#947 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
NVIDIA-NeMo/DataDesigner#946 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
task
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
NVIDIA-NeMo/DataDesigner#885 ·
Los mantenedores suelen responder en 1 día
Todos los issues de NVIDIA-NeMo/DataDesigner
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-2 días Aptitud para principiantes 70/100
-
FingerprintSplitter raises ZeroDivisionError when int(frac_train * len(dataset)) floors to zeroAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 7 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
lmstudio-ai/mlx-engine#376 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
pyiron/bagofholding#166 ·