[QNN] Enable ConvTranspose + BatchNorm fusion after #23170
Maintainer thường phản hồi trong vòng 1 ngày
@psiddh đang làm issue này rồi.
Từ ngày 28/9/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Feature and motivation
Enable ConvTranspose + BatchNorm fusion in the Qualcomm/QNN pass pipeline after #23170 lands.
FuseBatchNormWithConv.can_fuse currently rejects all transposed convolutions. This workaround avoids the incorrect output-channel scaling in the shared pass discussed in #22994. PR #23170 fixes that shared pass, including grouped transposed weights, but leaves the QNN guard in place. Consequently, QNN will continue to retain standalone BatchNorm operations after that fix.
Proposed work
- Remove the transposed-convolution guard once #23170 is merged, reusing the corrected shared implementation. Update the subclass documentation or remove the redundant wrapper as appropriate.
- Add Qualcomm regression cases for
ConvTranspose + BatchNorm, covering equal and unequal input/output channel counts, grouped convolutions, and convolution bias enabled/disabled where supported by QNN. - Use nontrivial BatchNorm running statistics and affine parameters. Assert that BatchNorm is folded and the transformed output matches eager execution.
- Run QNN delegation/execution coverage for the supported configurations and retain coverage for ordinary convolution fusion.
Additional context
Local CPU review validation of #23170 with PyTorch 2.13 passed all six new tests; four reproduce the original bug with the pre-fix pass. Eighteen additional grouped/depthwise ConvTranspose1d/2d/3d cases passed through the exported-program transform path. These checks validate the shared pass; QNN integration still needs the coverage above.
cc @cccclai @winskuo-quic @shewu-quic @haowhsu-quic @DannyYuyang-quic @cbilgin @abhinaykukkadapu @psiddh
- Ngôn ngữ chính
- Python
- Star
- 5k
- Fork
- 1.2k
- Merge trung bình
- 2 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 521
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của pytorch/executorch
-
enhancement triaged
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 68/100
pytorch/executorch#21640 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
pytorch/executorch#23263 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 68/100
pytorch/executorch#23262 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
XNNPACK never delegates slice_copy with stride != 1, but op-support.csv only excludes zero-dim/dynamic shapesCó thể đã có người làm @JakeStevens đã nhận hôm nay. Đang mởmodule: xnnpack
pytorch/executorch#23258 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[RFC] ExecuTorch Persisting Device Specialized Delegate ArtifactsCó thể đã có người làm @JacobSzwejbka đã nhận 1 ngày trước. Đang mở
pytorch/executorch#23192 · 1 bình luận · 1 reaction · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của pytorch/executorch
Issue tương tự
-
customer-reported
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Azure/azure-cli#34150 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
community-request
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 95/100
NVIDIA-NeMo/Curator#2464 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
weblate-discover crashes with an unhandled FileNotFoundError when the directory does not existĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
WeblateOrg/translation-finder#1099 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
trezor/trezor-firmware#7997 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày