Epic: Slurm batch execution v1
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- python
调研方向
先阅读基础 issue #851、#852 和 #853,然后继续查看通过依赖关系关联的契约,例如 #865、#873 和 #880。当可选 package 和 lazy CLI 已集成,并具备确定性规划、fake-runtime 覆盖、持久状态,以及列出的 Packaging、安全性、文档和 Slurm quality gate 时,此 epic 即完成。
由索引模型根据 Issue 内容生成。
描述
Priority Level
High
Task Summary
Track delivery of the first installable, batch-only Slurm integration for Data Designer.
The integration will ship as an optional data-designer-slurm package under the data_designer.slurm namespace. It will provide a strict configuration-to-result path that plans and submits Data Designer generation jobs, manages model services and client execution inside static Slurm allocations, and persists enough state for observation, retry, merge, and benchmark analysis.
This is a scoped delivery slice of #160. Interactive workflows and dynamic resource management remain future work.
Technical Details & Implementation Plan
Public outcome
- Install with
pip install "data-designer[slurm]"without adding Slurm dependencies to a base-only installation. - Expose an optional, lazily loaded
data-designer slurmcommand group and equivalent Python service APIs. - Provide
profile init/validate,execute,status,cancel,retry,merge,benchmark run/analyze, andimage add/ls/rm/infounder that command group. - Accept strict, versioned run, cluster-profile, serving, image, and benchmark inputs.
- Resolve authored input into an immutable execution plan before submission.
- Start model services and a separate zero-GPU Data Designer client within one static allocation per shard attempt.
- Persist run, shard, attempt, result, and output records so status, cancellation, retry, merge, and benchmark analysis work from a fresh process.
- Support deterministic local and fake-runtime tests before real Slurm validation.
Delivery stages
- Define the public Data Designer boundary and establish optional packaging and CLI discovery.
- Freeze only the schemas shared across packages, processes, or implementation lanes, integrate the plan and runtime-state boundaries, then add reusable fake infrastructure and golden fixtures.
- Implement configuration and planning, serving, image management, Slurm control/runtime, the Data Designer client worker, persistent state, and benchmarks in parallel where their contracts allow it.
- Validate local planning, image creation, one-node generation, multi-node serving, multiple serving images, sharding/retry/merge, and benchmarks.
- Complete security, packaging, documentation, and release checks against the exact artifacts selected for release.
Initial foundation work
- #851 Define the public Data Designer contract used by the optional Slurm package.
- #852 Add the optional
data-designer-slurmpackage and[slurm]extra. - #853 Add lazy discovery for optional CLI command groups.
- #865 Define shared runtime and state record contracts.
- #873 Define shared authored configuration, plan, image, client-result, and benchmark records.
- #880 Integrate Slurm plan and runtime-state contracts.
- #872 Add deterministic fake infrastructure for Slurm integration tests.
Later implementation lanes
- Public services, CLI commands, packaging, and documentation.
- Configuration resolution and deterministic plan compilation.
- #866 Typed serving resolution with vLLM support.
- #867 Container image import, inspection, and registry operations for the Slurm container runtime.
- #868 Slurm submission and allocation runtime.
- Data Designer client environment and worker.
- #869 Persistent state, retry, shard winner publication, and merge.
- Benchmark expansion, execution, observation, and analysis.
- #870 Runtime security review and sealed-artifact acceptance.
Out of scope for v1
- Interactive sessions, notebooks, tunnels, or detached model-server leases.
- Dynamic allocation resizing, early GPU release, or telemetry-driven resource changes.
- Kubernetes, Ray, a generic scheduler interface, or third-party serving plugins.
- Object-store output semantics.
- Runtime source checkouts, editable installs, or mutable dependency resolution.
Epic quality gates
- A base-only installation does not install or expose the Slurm package.
- The optional package uses only public Data Designer APIs and does not import engine internals.
- CLI extension discovery reads package metadata without importing the optional package until its command is selected.
- Authored inputs and persisted records are strict, versioned, deterministic, and redact secret values.
- Cluster-specific accounts, partitions, paths, mounts, and hardware facts are supplied through configuration rather than source branches or hardcoded defaults.
- Plans, runtime resources, images, state transitions, and published outputs are validated with checksums and atomic publication where applicable.
- Failed or partial attempts cannot become successful shard outputs, and merge selects only validated shard winners.
- Built-wheel installation tests, local/fake integration tests, real Slurm acceptance, documentation checks, dependency review, and source/package scans pass before release.
Investigation / Context
Related roadmap issue: #160.
The repository already uses plan and task issues to track large implementation efforts through dependency-ordered child issues. This epic follows that pattern while keeping the public scope limited to the new Data Designer capability.
Agent Plan / Findings
Start with the three foundation issues for the public Data Designer contract, optional package, and lazy CLI extension. Add issue links to the foundation checklist after those issues are created.
Shared-contract, fake-infrastructure, and feature-lane issues use public, standalone context. The plan/runtime-state boundary is tracked in #880 because it gates multiple downstream lanes; later environment-acceptance milestones remain quality gates on this epic rather than separate implementation issues.
Dependencies
Related to #160. Foundation work can start immediately.
- 主要语言
- Python
- 星标
- 2.3k
- 派生
- 215
- 平均合并
- 3 天 15 小时
- 30 天内合并 PR
- 41
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA-NeMo/DataDesigner 的其他 Issue
-
task
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA-NeMo/DataDesigner#760 ·
维护者通常 1 天内回复
-
bug
难度 3/5 1-2 天 新手友好度 78/100
NVIDIA-NeMo/DataDesigner#971 ·
维护者通常 1 天内回复
-
Run Slurm generation without a client image using a versioned vLLM serving image可能已有人在做 @nabinchha 今天认领。 未关闭task triaged
NVIDIA-NeMo/DataDesigner#969 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Harden Slurm inference routing, backpressure, and failover可能已有人在做 @nabinchha 于 1 天前认领。 未关闭task
NVIDIA-NeMo/DataDesigner#966 · 已指派 1 人 ·
维护者通常 1 天内回复
-
enhancement triaged
难度 4/5 3-5 天 新手友好度 40/100
NVIDIA-NeMo/DataDesigner#956 ·
维护者通常 1 天内回复
查看 NVIDIA-NeMo/DataDesigner 的全部 Issue
相似的 Issue
-
repo-audit
难度 2/5 1-3 小时 新手友好度 75/100
scverse/repo-health#20 ·
维护者通常 1 天内回复
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memories未关闭
难度 2/5 1-3 小时 新手友好度 85/100
phasespace-labs/palinode#232 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
collective/icalendar#1858 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
langflow-ai/langflow#15496 ·
维护者通常 1 天内回复