Qualcomm: 16a16w quantization is much less accurate than 16a8w
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 55/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
调研方向
Start by running the provided host-only reproduction and trace make_quantizer through get_16a16w_qnn_ptq_config. Compare the int16 per-tensor and per-channel weight specifications, including PerChannelParamObserver, to isolate the quantization-parameter discrepancy. Done means 16a16w no longer produces catastrophic SQNR or task-accuracy loss relative to 16a8w.
由索引模型根据 Issue 内容生成。
描述
QuantDtype.use_16a16w is far less accurate than use_16a8w on the same model. More weight bits should never mean a worse result.
resnet18, host PTQ only (prepare_pt2e / convert_pt2e, no lowering, no device), SQNR of the quantized output against fp32:
| recipe | SQNR vs fp32 |
|---|---|
use_16a8w |
24.49 dB |
use_16a16w |
-1.53 dB |
Same pattern on every model tried: deeplabv3_mobilenet_v3_large 8.14 -> 1.47 dB, mobilenet_v3_large 2.72 -> -0.75 dB. On a real task it is catastrophic: DeepLabV3-MobileNetV3 on VOC2012 scores 63.89% mIoU at 16a8w and 3.38% at 16a16w (fp32 67.52%, 100 calibration images).
Repro:
import torch
import torchvision.models as tvm
from executorch.backends.qualcomm.export_utils import make_quantizer
from executorch.backends.qualcomm.quantizer.quantizer import QuantDtype
from torchao.quantization.pt2e.quantize_pt2e import convert_pt2e, prepare_pt2e
torch.manual_seed(0)
model = tvm.resnet18(weights=tvm.ResNet18_Weights.DEFAULT).eval()
sample = (torch.randn(1, 3, 224, 224),)
probe = torch.randn(1, 3, 224, 224)
with torch.no_grad():
ref = model(probe)
for dtype in (QuantDtype.use_16a8w, QuantDtype.use_16a16w):
quantizer = make_quantizer(quant_dtype=dtype, soc_model="SM8850")
prepared = prepare_pt2e(
torch.export.export(model, sample, strict=False).module(), quantizer)
with torch.no_grad():
for _ in range(16):
prepared(torch.randn(1, 3, 224, 224))
got = convert_pt2e(prepared)(probe)
sqnr = 10 * torch.log10((ref**2).mean() / ((ref - got) ** 2).mean()).item()
print(f"{dtype.name:16s} SQNR vs fp32: {sqnr:7.2f} dB")
Ruled out:
- Not the int32 bias. Counting folded bias elements sitting at the int32 limits gives 0 saturated in both arms (DeepLabV3, 100 calibration images, 520px).
quantize_per_tensorpast int32 saturates rather than wraps, so this is not silent corruption there either. Worth noting separately that the derived bias scale leaves little headroom at 16-bit weights:max|bias/scale|reached 0.999 of int32 in one configuration. - Not the weights. Dequantized weights match between the two arms at 36-48 dB, i.e. the 16-bit weights are essentially exact.
- Not execution. This is fake-quant in PyTorch with convolutions in float, so no HTP and no integer accumulator are involved; the loss comes from the quantization parameters alone.
- Not an unsupported SoC.
validate_16a16w_supportrequires HTP >= V73 and SM8850 is v81.
The activation specs are identical between the two recipes, so the difference is on the weight side: get_16a16w_qnn_ptq_config uses an int16 per-tensor symmetric weight spec (quant_min=-32767, quant_max=32767), and the per-channel path uses weight_dtype=torch.int16 with PerChannelParamObserver. I have not isolated which of those is at fault.
Environment: executorch at 8081eb88, torch 2.14.0+cpu, torchao 0.18.0.dev20260729, torchvision 0.29.0+cpu, QAIRT 2.37.0.250724, Linux x86_64. No device needed.
cc @kimishpatel @jerryzh168 @metascroy @cbilgin @psiddh
- 主要语言
- Python
- 星标
- 5k
- 派生
- 1.2k
- 平均合并
- 2 天 9 小时
- 30 天内合并 PR
- 555
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
pytorch/executorch 的其他 Issue
-
enhancement triaged
难度 2/5 半天 新手友好度 68/100
pytorch/executorch#21640 ·
维护者通常 1 天内回复
-
enhancement module: examples
难度 5/5 一周以上 新手友好度 20/100
pytorch/executorch#23164 · 7 条评论 · 1 个 reaction ·
维护者通常 1 天内回复
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channels可能已有人在做 @psiddh 于 2 天前认领。 未关闭module: qnn partner: qualcomm
pytorch/executorch#23160 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[cpu kernels] native_layer_norm: layer_norm_scalar returns NaN on large-mean rows; Half/BF16 at N>=256 slow after #23153可能已有人在做 @JakeStevens 于 2 天前认领。 未关闭module: kernels
pytorch/executorch#23159 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
module: vulkan
难度 3/5 1-2 天 新手友好度 66/100
pytorch/executorch#23158 ·
维护者通常 1 天内回复
查看 pytorch/executorch 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 76/100
PedestrianDynamics/pyFDS-Evac#199 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
521xueweihan/HelloGitHub#3790 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
sandialabs/atlas-ui-3#978 ·
维护者通常 1 天内回复
-
area: tests perceived difficulty: 2
难度 2/5 1-3 小时 新手友好度 72/100
Nitjsefnie-Harness-Commons/daedalus#1255 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 86/100
EleutherAI/lm-evaluation-harness#4256 ·
维护者通常 1 天内回复