Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Nonfused column LoRA misses input-gradient SUM when TP>1 and sequence parallelism is disabled

オープン
#1,091 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

@bradhilton がすでに取り組んでいます。

2026年10月3日 から。

  • #1092 @bradhilton による — オープン

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python

調査の方向性

Start with src/art/megatron/lora.py, especially _column_parallel_lora_input and its SharedExpertsLinearFC1LoRA caller, then read the cited REPORT.md. Qualify a plain TE column projection with restored nonzero shared gate/up adapters using dense-only, adapter-only, and combined arms in both SP modes. Done means input VJPs and raw and synchronized parameter gradients match independent native references without duplicating outer-overlap reduction.

索引モデルが issue の本文から書いたものです。

説明

Agent: Schulman

A source audit found a missing input-gradient SUM in ordinary nonfused column-parallel LoRA when tensor parallelism is greater than one and sequence parallelism is disabled. The base TE projection reduces its own input gradient; the external adapter branch receives the input unchanged, so it contributes only the local shard’s cotangent. Later synchronization of adapter parameter gradients cannot repair that upstream input gradient.

Scope: nonzero gate/up adapters in SharedExpertsLinearFC1LoRA over plain TEColumnParallelLinear, with shared-expert overlap disabled. The same helper is used by the nonfused componentwise wrapper. Zero B or inactive adapters can hide the defect. The SP-enabled path already has gather/SUM-reduce-scatter, and externally owned shared overlap has a separate outer reduction; neither should acquire a duplicate reduction.

Relevant public source at reviewed #1087 head 2c3c929e773445c669838556f460ea0e8f9cd33e: _column_parallel_lora_input in src/art/megatron/lora.py returns the input unchanged for SP-disabled execution; SharedExpertsLinearFC1LoRA calls it for the nonfused branch. This path is unchanged by #1087’s fused-normalization correction.

Applicability limits: Qwen3.6 MoE shared FC1 reaches the ordinary nonfused wrapper, but the normal ART provider forces SP on when TP>1. The demonstrated exposure is a manually configured TP2/SP-disabled path, not an established failure of ordinary deployed Qwen TP2/SP-enabled runs. Unrestricted two-GPU Qwen defaults also use TP1/CP2/EP2. Do not attribute current packing errors to this finding.

Evidence: five actual-helper AST path controls and an exact two-rank algebra example distinguish the missing adapter SUM from double-reducing the base gradient. No native nonfused-wrapper outcome yet. Full source/caller/synchronization audit: /var/tmp/schulman-tp2-nonfused-input-source-sol61-20261003/REPORT.md, SHA256 0a079d7f6440745bbe62b1dda1f7840d8cd9645445570fe4068fcdccb10ca893.

Next qualification: actual plain TE column projection plus restored nonzero shared gate/up adapters, three dense-only/adapter-only/combined arms in both SP modes; compare input VJPs and raw/synchronized parameter gradients against independent native references. Preserve outer-overlap ownership. No tolerance or blanket correctness claim is established by the source/algebra checks. Owner: Schulman; queued behind current fused-wrapper and planner-memory qualification.

主要言語
Python
スター
10.8k
フォーク
1k
平均マージ
15時間 31分
マージ済み PR(30日)
166

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

OpenPipe/ART のほかの issue

OpenPipe/ART の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。