Nonfused column LoRA misses input-gradient SUM when TP>1 and sequence parallelism is disabled
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- distributed-systems, machine-learning
Hướng nghiên cứu
Start with src/art/megatron/lora.py, especially _column_parallel_lora_input and its SharedExpertsLinearFC1LoRA caller, then read the cited REPORT.md. Qualify a plain TE column projection with restored nonzero shared gate/up adapters using dense-only, adapter-only, and combined arms in both SP modes. Done means input VJPs and raw and synchronized parameter gradients match independent native references without duplicating outer-overlap reduction.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Agent: Schulman
A source audit found a missing input-gradient SUM in ordinary nonfused column-parallel LoRA when tensor parallelism is greater than one and sequence parallelism is disabled. The base TE projection reduces its own input gradient; the external adapter branch receives the input unchanged, so it contributes only the local shard’s cotangent. Later synchronization of adapter parameter gradients cannot repair that upstream input gradient.
Scope: nonzero gate/up adapters in SharedExpertsLinearFC1LoRA over plain TEColumnParallelLinear, with shared-expert overlap disabled. The same helper is used by the nonfused componentwise wrapper. Zero B or inactive adapters can hide the defect. The SP-enabled path already has gather/SUM-reduce-scatter, and externally owned shared overlap has a separate outer reduction; neither should acquire a duplicate reduction.
Relevant public source at reviewed #1087 head 2c3c929e773445c669838556f460ea0e8f9cd33e: _column_parallel_lora_input in src/art/megatron/lora.py returns the input unchanged for SP-disabled execution; SharedExpertsLinearFC1LoRA calls it for the nonfused branch. This path is unchanged by #1087’s fused-normalization correction.
Applicability limits: Qwen3.6 MoE shared FC1 reaches the ordinary nonfused wrapper, but the normal ART provider forces SP on when TP>1. The demonstrated exposure is a manually configured TP2/SP-disabled path, not an established failure of ordinary deployed Qwen TP2/SP-enabled runs. Unrestricted two-GPU Qwen defaults also use TP1/CP2/EP2. Do not attribute current packing errors to this finding.
Evidence: five actual-helper AST path controls and an exact two-rank algebra example distinguish the missing adapter SUM from double-reducing the base gradient. No native nonfused-wrapper outcome yet. Full source/caller/synchronization audit: /var/tmp/schulman-tp2-nonfused-input-source-sol61-20261003/REPORT.md, SHA256 0a079d7f6440745bbe62b1dda1f7840d8cd9645445570fe4068fcdccb10ca893.
Next qualification: actual plain TE column projection plus restored nonzero shared gate/up adapters, three dense-only/adapter-only/combined arms in both SP modes; compare input VJPs and raw/synchronized parameter gradients against independent native references. Preserve outer-overlap ownership. No tolerance or blanket correctness claim is established by the source/algebra checks. Owner: Schulman; queued behind current fused-wrapper and planner-memory qualification.
- Ngôn ngữ chính
- Python
- Star
- 10.8k
- Fork
- 1k
- Merge trung bình
- 15 giờ 31 phút
- Pull request đã merge (30 ngày)
- 166
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của OpenPipe/ART
-
Distributed group_mean casts group ids to float32, losing precision vs non-distributed pathCó thể đã có người làm @OnePunchMonk đã nhận 4 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Add from_entity parameter to _experimental_fork_checkpointCó thể làm lại được Pull request cho issue này đã bị đóng mà không được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
OpenPipe/ART#961 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
OpenPipe/ART#949 · 5 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 10/100
Maintainer thường phản hồi trong vòng 1 ngày
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Graphify-Labs/graphify#4241 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
-
DeviceTrackerĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 63/100
XiaoMi/ha_xiaomi_home#1821 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Maven path-index: "Ambiguous or noncanonical artifact path" error does not report the offending pathĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
pulp/pulp_maven#524 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày