Support pipeline parallelism in TrainerRank without exposing stage scheduling to callers
Maintainer thường phản hồi trong vòng 3 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- distributed-systems, machine-learning
Hướng nghiên cứu
Bắt đầu trong src/art/trainer_rank/_impl.py tại guard của constructor được tham chiếu trong issue, sau đó lần theo các đường dẫn public của forward, backward và optimizer. Xác định và triển khai việc lập lịch pipeline và giao tiếp nội bộ, đồng thời giữ nguyên ngữ nghĩa API được liệt kê. Được xem là hoàn tất khi canary PP=2 và các phép so sánh PP=1 đều vượt qua, bao gồm no-grad, tích lũy, các microbatch không đều, lựa chọn checkpoint và các lỗi tường minh đối với những tổ hợp không được hỗ trợ.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
TrainerRank currently rejects pipeline parallelism (PP > 1) and multiple local model chunks. Support pipeline-parallel execution internally so experiment code can keep using the same forwards, full outputs, custom losses, backward, and optimizer API without managing pipeline stages.
Current behavior
At ART main 7496cc09252c52ec7a63ab74d11caabce173f5b0, the constructor raises TrainerRankRuntimeSupportError when runtime.provider.pipeline_model_parallel_size > 1 or len(runtime.model) > 1. The error explains that TrainerRank does not use the MCore forward/backward schedule and requires PP=1 with exactly one local model chunk.
This is a source-confirmed unsupported configuration, not a newly observed GPU failure. Removing the guard alone would not provide the missing scheduling and communication.
Desired behavior
- Internally schedule stage forwards/backwards and activation/gradient transfers while preserving
forward_micro_batchesanddp_rank_forwardsemantics. - Preserve full source-order outputs, caller-defined losses and registered custom heads, checkpoint selection, gradient accumulation, and optimizer behavior. Public DP reductions must not count pipeline stages as independent data batches.
- Keep pipeline stage ownership and scheduling out of experiment code. Explicitly define supported combinations with TP/CP and any initial limitations, including multiple local chunks.
Acceptance
- A native PP=2 canary completes forward, caller-side loss, backward, and optimizer update through the public API.
- Compare outputs, model/custom-head gradients, and parameter updates with a matched PP=1 reference within stated numerical tolerances; cover no-grad execution and accumulation across multiple forwards.
- Exercise uneven microbatch workloads and checkpoint selection without mismatched communication or hangs. Keep explicit errors for configurations that remain unsupported.
Related: #911 / #912 establish the full-output contract for context parallelism; this issue tracks the separate pipeline execution limitation.
- Ngôn ngữ chính
- Python
- Star
- 10.8k
- Fork
- 989
- Merge trung bình
- 11 giờ 38 phút
- Pull request đã merge (30 ngày)
- 104
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của OpenPipe/ART
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
OpenPipe/ART#961 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
OpenPipe/ART#949 · 5 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 10/100
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 42/100
Maintainer thường phản hồi trong vòng 3 ngày
Issue tương tự
-
Broken links found in docsĐang mởdocs pydanty:is-working
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
pydantic/pydantic-ai#8863 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
run-llama/llama_index#23278 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
documentation from-review-extraction github-actions priority: low severity:nit
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
LearningCircuit/local-deep-research#6946 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
oracle/langchain-oracle#323 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
tenstorrent/tt-metal#58057 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày