Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Epic: Slurm batch execution v1

未关闭
#850 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
25/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
冷清
技术栈
python

调研方向

先阅读基础 issue #851、#852 和 #853,然后继续查看通过依赖关系关联的契约,例如 #865、#873 和 #880。当可选 package 和 lazy CLI 已集成,并具备确定性规划、fake-runtime 覆盖、持久状态,以及列出的 Packaging、安全性、文档和 Slurm quality gate 时,此 epic 即完成。

由索引模型根据 Issue 内容生成。

描述

plan task
Priority Level

High

Task Summary

Track delivery of the first installable, batch-only Slurm integration for Data Designer.

The integration will ship as an optional data-designer-slurm package under the data_designer.slurm namespace. It will provide a strict configuration-to-result path that plans and submits Data Designer generation jobs, manages model services and client execution inside static Slurm allocations, and persists enough state for observation, retry, merge, and benchmark analysis.

This is a scoped delivery slice of #160. Interactive workflows and dynamic resource management remain future work.

Technical Details & Implementation Plan
Public outcome
  • Install with pip install "data-designer[slurm]" without adding Slurm dependencies to a base-only installation.
  • Expose an optional, lazily loaded data-designer slurm command group and equivalent Python service APIs.
  • Provide profile init/validate, execute, status, cancel, retry, merge, benchmark run/analyze, and image add/ls/rm/info under that command group.
  • Accept strict, versioned run, cluster-profile, serving, image, and benchmark inputs.
  • Resolve authored input into an immutable execution plan before submission.
  • Start model services and a separate zero-GPU Data Designer client within one static allocation per shard attempt.
  • Persist run, shard, attempt, result, and output records so status, cancellation, retry, merge, and benchmark analysis work from a fresh process.
  • Support deterministic local and fake-runtime tests before real Slurm validation.
Delivery stages
  1. Define the public Data Designer boundary and establish optional packaging and CLI discovery.
  2. Freeze only the schemas shared across packages, processes, or implementation lanes, integrate the plan and runtime-state boundaries, then add reusable fake infrastructure and golden fixtures.
  3. Implement configuration and planning, serving, image management, Slurm control/runtime, the Data Designer client worker, persistent state, and benchmarks in parallel where their contracts allow it.
  4. Validate local planning, image creation, one-node generation, multi-node serving, multiple serving images, sharding/retry/merge, and benchmarks.
  5. Complete security, packaging, documentation, and release checks against the exact artifacts selected for release.
Initial foundation work
  • #851 Define the public Data Designer contract used by the optional Slurm package.
  • #852 Add the optional data-designer-slurm package and [slurm] extra.
  • #853 Add lazy discovery for optional CLI command groups.
  • #865 Define shared runtime and state record contracts.
  • #873 Define shared authored configuration, plan, image, client-result, and benchmark records.
  • #880 Integrate Slurm plan and runtime-state contracts.
  • #872 Add deterministic fake infrastructure for Slurm integration tests.
Later implementation lanes
  • Public services, CLI commands, packaging, and documentation.
  • Configuration resolution and deterministic plan compilation.
  • #866 Typed serving resolution with vLLM support.
  • #867 Container image import, inspection, and registry operations for the Slurm container runtime.
  • #868 Slurm submission and allocation runtime.
  • Data Designer client environment and worker.
  • #869 Persistent state, retry, shard winner publication, and merge.
  • Benchmark expansion, execution, observation, and analysis.
  • #870 Runtime security review and sealed-artifact acceptance.
Out of scope for v1
  • Interactive sessions, notebooks, tunnels, or detached model-server leases.
  • Dynamic allocation resizing, early GPU release, or telemetry-driven resource changes.
  • Kubernetes, Ray, a generic scheduler interface, or third-party serving plugins.
  • Object-store output semantics.
  • Runtime source checkouts, editable installs, or mutable dependency resolution.
Epic quality gates
  • A base-only installation does not install or expose the Slurm package.
  • The optional package uses only public Data Designer APIs and does not import engine internals.
  • CLI extension discovery reads package metadata without importing the optional package until its command is selected.
  • Authored inputs and persisted records are strict, versioned, deterministic, and redact secret values.
  • Cluster-specific accounts, partitions, paths, mounts, and hardware facts are supplied through configuration rather than source branches or hardcoded defaults.
  • Plans, runtime resources, images, state transitions, and published outputs are validated with checksums and atomic publication where applicable.
  • Failed or partial attempts cannot become successful shard outputs, and merge selects only validated shard winners.
  • Built-wheel installation tests, local/fake integration tests, real Slurm acceptance, documentation checks, dependency review, and source/package scans pass before release.
Investigation / Context

Related roadmap issue: #160.

The repository already uses plan and task issues to track large implementation efforts through dependency-ordered child issues. This epic follows that pattern while keeping the public scope limited to the new Data Designer capability.

Agent Plan / Findings

Start with the three foundation issues for the public Data Designer contract, optional package, and lazy CLI extension. Add issue links to the foundation checklist after those issues are created.

Shared-contract, fake-infrastructure, and feature-lane issues use public, standalone context. The plan/runtime-state boundary is tracked in #880 because it gates multiple downstream lanes; later environment-acceptance milestones remain quality gates on this epic rather than separate implementation issues.

Dependencies

Related to #160. Foundation work can start immediately.

主要语言
Python
星标
2.3k
派生
215
平均合并
3 天 15 小时
30 天内合并 PR
41

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA-NeMo/DataDesigner 的其他 Issue

查看 NVIDIA-NeMo/DataDesigner 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。