Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Support pipeline parallelism in TrainerRank without exposing stage scheduling to callers

Đang mở
#916 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 3 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python

Hướng nghiên cứu

Bắt đầu trong src/art/trainer_rank/_impl.py tại guard của constructor được tham chiếu trong issue, sau đó lần theo các đường dẫn public của forward, backward và optimizer. Xác định và triển khai việc lập lịch pipeline và giao tiếp nội bộ, đồng thời giữ nguyên ngữ nghĩa API được liệt kê. Được xem là hoàn tất khi canary PP=2 và các phép so sánh PP=1 đều vượt qua, bao gồm no-grad, tích lũy, các microbatch không đều, lựa chọn checkpoint và các lỗi tường minh đối với những tổ hợp không được hỗ trợ.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

TrainerRank currently rejects pipeline parallelism (PP > 1) and multiple local model chunks. Support pipeline-parallel execution internally so experiment code can keep using the same forwards, full outputs, custom losses, backward, and optimizer API without managing pipeline stages.

Current behavior

At ART main 7496cc09252c52ec7a63ab74d11caabce173f5b0, the constructor raises TrainerRankRuntimeSupportError when runtime.provider.pipeline_model_parallel_size > 1 or len(runtime.model) > 1. The error explains that TrainerRank does not use the MCore forward/backward schedule and requires PP=1 with exactly one local model chunk.

This is a source-confirmed unsupported configuration, not a newly observed GPU failure. Removing the guard alone would not provide the missing scheduling and communication.

Desired behavior
  • Internally schedule stage forwards/backwards and activation/gradient transfers while preserving forward_micro_batches and dp_rank_forward semantics.
  • Preserve full source-order outputs, caller-defined losses and registered custom heads, checkpoint selection, gradient accumulation, and optimizer behavior. Public DP reductions must not count pipeline stages as independent data batches.
  • Keep pipeline stage ownership and scheduling out of experiment code. Explicitly define supported combinations with TP/CP and any initial limitations, including multiple local chunks.
Acceptance
  • A native PP=2 canary completes forward, caller-side loss, backward, and optimizer update through the public API.
  • Compare outputs, model/custom-head gradients, and parameter updates with a matched PP=1 reference within stated numerical tolerances; cover no-grad execution and accumulation across multiple forwards.
  • Exercise uneven microbatch workloads and checkpoint selection without mismatched communication or hangs. Keep explicit errors for configurations that remain unsupported.

Related: #911 / #912 establish the full-output contract for context parallelism; this issue tracks the separate pipeline execution limitation.

Ngôn ngữ chính
Python
Star
10.8k
Fork
989
Merge trung bình
11 giờ 38 phút
Pull request đã merge (30 ngày)
104

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của OpenPipe/ART

Tất cả issue của OpenPipe/ART

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.