[QNN] Enable ConvTranspose + BatchNorm fusion after #23170
Les mainteneurs répondent en général sous 1 jour
@psiddh y travaille déjà.
Depuis le 28/9/2026.
Évaluation
Cette issue n'a pas encore été évaluée.
Description
Feature and motivation
Enable ConvTranspose + BatchNorm fusion in the Qualcomm/QNN pass pipeline after #23170 lands.
FuseBatchNormWithConv.can_fuse currently rejects all transposed convolutions. This workaround avoids the incorrect output-channel scaling in the shared pass discussed in #22994. PR #23170 fixes that shared pass, including grouped transposed weights, but leaves the QNN guard in place. Consequently, QNN will continue to retain standalone BatchNorm operations after that fix.
Proposed work
- Remove the transposed-convolution guard once #23170 is merged, reusing the corrected shared implementation. Update the subclass documentation or remove the redundant wrapper as appropriate.
- Add Qualcomm regression cases for
ConvTranspose + BatchNorm, covering equal and unequal input/output channel counts, grouped convolutions, and convolution bias enabled/disabled where supported by QNN. - Use nontrivial BatchNorm running statistics and affine parameters. Assert that BatchNorm is folded and the transformed output matches eager execution.
- Run QNN delegation/execution coverage for the supported configurations and retain coverage for ordinary convolution fusion.
Additional context
Local CPU review validation of #23170 with PyTorch 2.13 passed all six new tests; four reproduce the original bug with the pre-fix pass. Eighteen additional grouped/depthwise ConvTranspose1d/2d/3d cases passed through the exported-program transform path. These checks validate the shared pass; QNN integration still needs the coverage above.
cc @cccclai @winskuo-quic @shewu-quic @haowhsu-quic @DannyYuyang-quic @cbilgin @abhinaykukkadapu @psiddh
- Langage dominant
- Python
- Étoiles
- 5k
- Forks
- 1.2k
- Merge moyen
- 2 j 7 h
- PR mergées (30 j)
- 521
Préparer son environnement
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de pytorch/executorch
-
enhancement triaged
Difficulté 2/5 Une demi-journée Accessibilité débutants 68/100
pytorch/executorch#21640 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 4/5 3-5 jours Accessibilité débutants 55/100
pytorch/executorch#23263 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 4/5 3-5 jours Accessibilité débutants 68/100
pytorch/executorch#23262 ·
Les mainteneurs répondent en général sous 1 jour
-
XNNPACK never delegates slice_copy with stride != 1, but op-support.csv only excludes zero-dim/dynamic shapesPeut-être pris @JakeStevens l’a pris aujourd’hui. Ouvertemodule: xnnpack
pytorch/executorch#23258 · 1 personne assignée ·
Les mainteneurs répondent en général sous 1 jour
-
[RFC] ExecuTorch Persisting Device Specialized Delegate ArtifactsPeut-être pris @JacobSzwejbka l’a pris il y a 1 jour. Ouverte
pytorch/executorch#23192 · 1 commentaire · 1 réaction · 1 personne assignée ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de pytorch/executorch
Issues similaires
-
bug
Difficulté 2/5 1-3 heures Accessibilité débutants 85/100
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 1/5 Moins d'une heure Accessibilité débutants 90/100
Les mainteneurs répondent en général sous 1 jour
-
https://search.utilibre.orgOuverteinstance instance add
Difficulté 2/5 1-3 heures Accessibilité débutants 68/100
searxng/searx-instances#941 · 1 commentaire ·
-
Difficulté 1/5 Moins d'une heure Accessibilité débutants 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Les mainteneurs répondent en général sous 1 jour
-
bug
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
Les mainteneurs répondent en général sous 1 jour