Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Qualcomm: 16a16w quantization is much less accurate than 16a8w

未关闭
#23,108 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
55/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
python, pytorch

调研方向

Start by running the provided host-only reproduction and trace make_quantizer through get_16a16w_qnn_ptq_config. Compare the int16 per-tensor and per-channel weight specifications, including PerChannelParamObserver, to isolate the quantization-parameter discrepancy. Done means 16a16w no longer produces catastrophic SQNR or task-accuracy loss relative to 16a8w.

由索引模型根据 Issue 内容生成。

描述

bug module: qnn module: quantization

QuantDtype.use_16a16w is far less accurate than use_16a8w on the same model. More weight bits should never mean a worse result.

resnet18, host PTQ only (prepare_pt2e / convert_pt2e, no lowering, no device), SQNR of the quantized output against fp32:

recipe SQNR vs fp32
use_16a8w 24.49 dB
use_16a16w -1.53 dB

Same pattern on every model tried: deeplabv3_mobilenet_v3_large 8.14 -> 1.47 dB, mobilenet_v3_large 2.72 -> -0.75 dB. On a real task it is catastrophic: DeepLabV3-MobileNetV3 on VOC2012 scores 63.89% mIoU at 16a8w and 3.38% at 16a16w (fp32 67.52%, 100 calibration images).

Repro:

import torch
import torchvision.models as tvm
from executorch.backends.qualcomm.export_utils import make_quantizer
from executorch.backends.qualcomm.quantizer.quantizer import QuantDtype
from torchao.quantization.pt2e.quantize_pt2e import convert_pt2e, prepare_pt2e

torch.manual_seed(0)
model = tvm.resnet18(weights=tvm.ResNet18_Weights.DEFAULT).eval()
sample = (torch.randn(1, 3, 224, 224),)
probe = torch.randn(1, 3, 224, 224)
with torch.no_grad():
    ref = model(probe)

for dtype in (QuantDtype.use_16a8w, QuantDtype.use_16a16w):
    quantizer = make_quantizer(quant_dtype=dtype, soc_model="SM8850")
    prepared = prepare_pt2e(
        torch.export.export(model, sample, strict=False).module(), quantizer)
    with torch.no_grad():
        for _ in range(16):
            prepared(torch.randn(1, 3, 224, 224))
        got = convert_pt2e(prepared)(probe)
    sqnr = 10 * torch.log10((ref**2).mean() / ((ref - got) ** 2).mean()).item()
    print(f"{dtype.name:16s} SQNR vs fp32: {sqnr:7.2f} dB")

Ruled out:

  • Not the int32 bias. Counting folded bias elements sitting at the int32 limits gives 0 saturated in both arms (DeepLabV3, 100 calibration images, 520px). quantize_per_tensor past int32 saturates rather than wraps, so this is not silent corruption there either. Worth noting separately that the derived bias scale leaves little headroom at 16-bit weights: max|bias/scale| reached 0.999 of int32 in one configuration.
  • Not the weights. Dequantized weights match between the two arms at 36-48 dB, i.e. the 16-bit weights are essentially exact.
  • Not execution. This is fake-quant in PyTorch with convolutions in float, so no HTP and no integer accumulator are involved; the loss comes from the quantization parameters alone.
  • Not an unsupported SoC. validate_16a16w_support requires HTP >= V73 and SM8850 is v81.

The activation specs are identical between the two recipes, so the difference is on the weight side: get_16a16w_qnn_ptq_config uses an int16 per-tensor symmetric weight spec (quant_min=-32767, quant_max=32767), and the per-channel path uses weight_dtype=torch.int16 with PerChannelParamObserver. I have not isolated which of those is at fault.

Environment: executorch at 8081eb88, torch 2.14.0+cpu, torchao 0.18.0.dev20260729, torchvision 0.29.0+cpu, QAIRT 2.37.0.250724, Linux x86_64. No device needed.

cc @kimishpatel @jerryzh168 @metascroy @cbilgin @psiddh

主要语言
Python
星标
5k
派生
1.2k
平均合并
2 天 9 小时
30 天内合并 PR
555

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

pytorch/executorch 的其他 Issue

查看 pytorch/executorch 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。