[PyTorch][Attention] THD P2P context-parallel regression when padded cu_seqlens are value-equal but not object-identical
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
- #3506 de @RPalmr — cerrado sin fusionar
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 38/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
Línea de trabajo
Comienza por el detector de padding de PyTorch DotProductAttention introducido en el commit 4cd705b75394563c0246bdddfa5d3148106c9285; después, inspecciona get_cu_seqlens_on_cp_rank y la ruta de puesta a cero de la cola dQ/dK/dV de THD. Reproduce los casos alternos de pad_between_seqs con el benchmark proporcionado y compara el trabajo del kernel y los tiempos de ejecución. Se considera completado cuando los metadatos equivalentes en valor evitan trabajo de padding innecesario, mientras el comportamiento numérico y la operación segura para grafos permanecen intactos.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Describe the bug
Commit 4cd705b75394563c0246bdddfa5d3148106c9285 introduces a PyTorch DotProductAttention performance regression under this combination:
qkv_format="thd"- P2P context parallelism
pad_between_seqs=None(automatic detection)- both regular and padded cumulative sequence-length tensors are provided
- the padded and unpadded tensors are distinct objects but have identical relevant values, so there is no actual padding between sequences
This is observable during ordinary eager training; CUDA graph capture does not need to be enabled.
The new automatic detection uses tensor object identity as a proxy for padding semantics:
if cu_seqlens_q_padded is cu_seqlens_q:
pad_between_seqs = False
elif cu_seqlens_q_padded is not None or cu_seqlens_kv_padded is not None:
pad_between_seqs = True
Thus independently allocated but value-identical tensors are classified as pad_between_seqs=True. With P2P context parallelism, that selects the path that repeatedly calls get_cu_seqlens_on_cp_rank, rather than the cheaper no-inter-sequence-padding path.
The same commit also adds THD dQ/dK/dV tail-zeroing operations. They launch arange/compare/masked-fill work even when the valid endpoint already equals the tensor endpoint and there is no tail to clear.
Steps/Code to reproduce bug
-
Create a BF16
DotProductAttentionmodule withqkv_format="thd"and a four-rank P2P context-parallel group. -
Provide independently allocated cumulative sequence-length tensors with identical values:
cu_seqlens = torch.tensor([0, sequence_length], dtype=torch.int32, device="cuda") cu_seqlens_padded = cu_seqlens.clone() assert cu_seqlens_padded is not cu_seqlens assert torch.equal(cu_seqlens_padded, cu_seqlens) -
Run repeated attention forward/backward calls, alternating these two cases in the same process:
- automatic detection:
pad_between_seqs=None - known-correct metadata:
pad_between_seqs=False
- automatic detection:
-
Discard warmup and compare steady-state timings. A single attention forward/backward call shows a small direct overhead. The impact becomes much larger in an attention-heavy training schedule where the branch is exercised repeatedly and interacts with context-parallel stream scheduling.
We also performed a controlled source-level reverse experiment on Transformer Engine 2.18.0+27486e03. All arms used the same process, allocation, inputs, configuration, and byte-identical compiled Transformer Engine extensions; only the Python attention hunks from the cited commit differed. Each arm used 50 post-warmup iterations.
| Variant | Mean iteration time | Median | Delta vs. stock |
|---|---|---|---|
| Stock | 595.084 ms | 589.000 ms | — |
| Revert padding detector only | 560.958 ms | 556.800 ms | -5.735% |
| Revert gradient zero-fill only | 577.764 ms | 574.950 ms | -2.911% |
| Revert both | 554.254 ms | 549.650 ms | -6.861% |
An ABBA repetition of stock and the full reverse patch measured a 6.276% aggregate iteration-time improvement with the reverse patch. The stock and reverse-patched order drift was 1.534% and 0.861%, respectively.
Across all ranks in two Nsight Systems trials, the full reverse patch reduced the five-step trace span by 7.411% on average, with every paired rank faster. Over five steps it removed, per rank:
- 2,400 helper-generated kernel launches associated with
get_cu_seqlens_on_cp_rank - 900 masked-fill kernels from the new gradient tail-zeroing blocks
The source reversal reduced main-stream kernel work by 31.702 ms/rank and main-stream gaps by 237.154 ms/rank over the captured five-step window. These are separate trace observations, not additive wall-time attribution. Numerical-health checks remained clean.
Expected behavior
When the padded and unpadded cumulative sequence-length tensors have equal relevant values, automatic detection should not select the inter-sequence-padding path solely because they are different Python objects.
Could the API carry graph-safe padding metadata explicitly, or otherwise avoid using object identity as the semantic proxy? The tail-zeroing work could also be gated when metadata establishes that no gradient tail exists, while preserving CUDA graph compatibility.
The current workaround for callers that know there is no inter-sequence padding is to pass pad_between_seqs=False explicitly.
Environment overview
- Environment location: containerized bare-metal system
- Transformer Engine:
2.18.0+27486e03 - Installation: preinstalled container package
Environment details
- Python: 3.12.3
- PyTorch:
2.13.0a0+8145d630e8.nv26.6.54250401 - CUDA reported by PyTorch: 13.3
- cuDNN: compiled against 9.23; node-visible runtime 9.21.1
The causal comparison used one unchanged environment, so the cuDNN packaging detail was identical across all variants.
Device details
- 4x NVIDIA H100 80GB HBM3
Additional context
The detector is the primary contributor. After removing the zero-fill blocks, reverting the detector still improved iteration time by 4.069%. Once the detector was corrected, removing zero-fill added another 1.195%. The effects overlap on the same context-parallel critical path and therefore should not be added independently.
- Lenguaje dominante
- Python
- Estrellas
- 3.6k
- Forks
- 851
- Merge medio
- 5 d 1 h
- PR fusionados (30 d)
- 52
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA/TransformerEngine
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
NVIDIA/TransformerEngine#3647 ·
Los mantenedores suelen responder en 2 días
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarPosiblemente ocupada @sanjana658 la tomó hace 1 día. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
NVIDIA/TransformerEngine#3636 · 2 comentarios ·
Los mantenedores suelen responder en 2 días
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itPosiblemente ocupada @yuweih205 la tomó hace 32 días. Abiertoattention
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
NVIDIA/TransformerEngine#3481 · 4 comentarios ·
Los mantenedores suelen responder en 2 días
-
Increase MAX_TENSOR_NUMAbiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
NVIDIA/TransformerEngine#2189 · 7 comentarios · 5 reacciones ·
Los mantenedores suelen responder en 2 días
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration buildsPosiblemente ocupada @ksivaman la tomó hace 1 día. Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 50/100
NVIDIA/TransformerEngine#3645 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 2 días
Todos los issues de NVIDIA/TransformerEngine
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
MystenLabs/MemWal#1163 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
infertopics leaves new nodes without a topic when untopiced neighbours outnumber topiced onesPosiblemente ocupada @moneebullah25 la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
ClanGenOfficial/clangen#6254 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
FinanceFlash/unvibecode#218 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Los mantenedores suelen responder en 1 día