Resolve deployment-owned vLLM serving plans for Slurm
Maintainer thường phản hồi trong vòng 1 ngày
@nabinchha đang làm issue này rồi.
Từ ngày 15/8/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Priority Level
High
Task Summary
Implement the package-owned vLLM resolver that converts one typed deployment declaration plus planner-supplied placement into process, readiness, backend-endpoint, and logical-endpoint records.
Technical Details & Implementation Plan
- Dispatch on the deployment server discriminator with vLLM as the only implemented backend.
- Keep model alias, model source, and optional served model name distinct.
- Resolve tensor parallelism within a node and pipeline parallelism across nodes in a replica group.
- Derive replica groups, ranks, lanes, rendezvous inputs, and lane-head backend endpoints from planner-owned placement.
- Carry per-deployment startup and distributed-initialization deadlines, launch standoff/stagger, queue backpressure, readiness configuration, tokenized extra arguments, and typed environment bindings.
- Reject extra arguments that override compiler/runtime-owned model identity, ports, topology, rendezvous, launch timing, or middleware behavior.
- Emit typed logical-endpoint inputs so the runtime can aggregate readiness and load-balance healthy backends while preserving overload responses.
- Preserve the inspected vLLM runtime version as digest-bound provenance without maintaining a package-version compatibility matrix.
Acceptance criteria
- A single-node deployment resolves deterministically without importing scheduler or shell modules.
- A multi-node deployment derives the expected replica groups, ranks, pipeline parallelism, and lane-head endpoints.
- Two deployments may select different serving images without endpoint or resource identity collisions.
- Invalid node, GPU, tensor-parallel, replica-group, expert-parallel, image-kind, and resolved-behavior combinations fail before submission.
- Default and overridden launch-timing and queue-backpressure values serialize in golden records.
- A rank failure is represented as failure of the coordinated deployment; no per-replica restart contract is introduced.
Out of scope
- Slurm submission or node allocation.
- Shell rendering, process supervision, and cleanup.
- Image building or registry mutation.
- Dynamo, SGLang, serving plugins, and per-replica recovery.
Investigation / Context
This is a feature lane under #850. Portable deployment intent remains separate from backend-specific process resolution so configuration stays declarative.
Agent Plan / Findings
Build against the reviewed shared configuration, image-inspection, placement, and runtime record contracts. Keep the resolver pure with focused single-node and multi-node golden tests.
Dependencies
Blocked by the shared-contract and fake-infrastructure foundation work tracked by #850.
- Ngôn ngữ chính
- Python
- Star
- 2.3k
- Fork
- 215
- Merge trung bình
- 3 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 45
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA-NeMo/DataDesigner
-
task
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 78/100
NVIDIA-NeMo/DataDesigner#971 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Run Slurm generation without a client image using a versioned vLLM serving imageCó thể đã có người làm @nabinchha đã nhận hôm nay. Đang mởtask triaged
NVIDIA-NeMo/DataDesigner#969 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Harden Slurm inference routing, backpressure, and failoverCó thể đã có người làm @nabinchha đã nhận 1 ngày trước. Đang mởtask
NVIDIA-NeMo/DataDesigner#966 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement triaged
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 40/100
NVIDIA-NeMo/DataDesigner#956 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của NVIDIA-NeMo/DataDesigner
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
kornia/kornia#5263 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Metadata correction for W16-5400Đang mởapproved correction metadata
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
acl-org/acl-anthology#10133 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
BasedHardware/omi#20084 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug needs-acceptance wg/evaluation-quality
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
vllm-project/semantic-router#4424 ·
Maintainer thường phản hồi trong vòng 1 ngày