Nonfused column LoRA misses input-gradient SUM when TP>1 and sequence parallelism is disabled
Los mantenedores suelen responder en 1 día
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
Línea de trabajo
Start with src/art/megatron/lora.py, especially _column_parallel_lora_input and its SharedExpertsLinearFC1LoRA caller, then read the cited REPORT.md. Qualify a plain TE column projection with restored nonzero shared gate/up adapters using dense-only, adapter-only, and combined arms in both SP modes. Done means input VJPs and raw and synchronized parameter gradients match independent native references without duplicating outer-overlap reduction.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Agent: Schulman
A source audit found a missing input-gradient SUM in ordinary nonfused column-parallel LoRA when tensor parallelism is greater than one and sequence parallelism is disabled. The base TE projection reduces its own input gradient; the external adapter branch receives the input unchanged, so it contributes only the local shard’s cotangent. Later synchronization of adapter parameter gradients cannot repair that upstream input gradient.
Scope: nonzero gate/up adapters in SharedExpertsLinearFC1LoRA over plain TEColumnParallelLinear, with shared-expert overlap disabled. The same helper is used by the nonfused componentwise wrapper. Zero B or inactive adapters can hide the defect. The SP-enabled path already has gather/SUM-reduce-scatter, and externally owned shared overlap has a separate outer reduction; neither should acquire a duplicate reduction.
Relevant public source at reviewed #1087 head 2c3c929e773445c669838556f460ea0e8f9cd33e: _column_parallel_lora_input in src/art/megatron/lora.py returns the input unchanged for SP-disabled execution; SharedExpertsLinearFC1LoRA calls it for the nonfused branch. This path is unchanged by #1087’s fused-normalization correction.
Applicability limits: Qwen3.6 MoE shared FC1 reaches the ordinary nonfused wrapper, but the normal ART provider forces SP on when TP>1. The demonstrated exposure is a manually configured TP2/SP-disabled path, not an established failure of ordinary deployed Qwen TP2/SP-enabled runs. Unrestricted two-GPU Qwen defaults also use TP1/CP2/EP2. Do not attribute current packing errors to this finding.
Evidence: five actual-helper AST path controls and an exact two-rank algebra example distinguish the missing adapter SUM from double-reducing the base gradient. No native nonfused-wrapper outcome yet. Full source/caller/synchronization audit: /var/tmp/schulman-tp2-nonfused-input-source-sol61-20261003/REPORT.md, SHA256 0a079d7f6440745bbe62b1dda1f7840d8cd9645445570fe4068fcdccb10ca893.
Next qualification: actual plain TE column projection plus restored nonzero shared gate/up adapters, three dense-only/adapter-only/combined arms in both SP modes; compare input VJPs and raw/synchronized parameter gradients against independent native references. Preserve outer-overlap ownership. No tolerance or blanket correctness claim is established by the source/algebra checks. Owner: Schulman; queued behind current fused-wrapper and planner-memory qualification.
- Lenguaje dominante
- Python
- Estrellas
- 10.8k
- Forks
- 1k
- Merge medio
- 15 h 31 min
- PR fusionados (30 d)
- 166
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de OpenPipe/ART
-
Distributed group_mean casts group ids to float32, losing precision vs non-distributed pathPosiblemente ocupada @OnePunchMonk la tomó hace 3 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
Add from_entity parameter to _experimental_fork_checkpointQuizá libre de nuevo Un pull request para esta issue se cerró sin fusionarse. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 54/100
OpenPipe/ART#961 · 3 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
OpenPipe/ART#949 · 5 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 10/100
Los mantenedores suelen responder en 1 día
Todos los issues de OpenPipe/ART
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
-
Harmony OPeNDAP SubSetter (HOSS) Geographic LARC_CLOUD PREFIRE_SAT2_AUX-SAT R01 production
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
nasa/harmony-autotester#245 ·
-
[FEATURE] - Add UTVD supportAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Deltares/imod-python#1928 ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
feature
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100