Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[winml perf] Compiled w8a16 model is incorrectly reported as w16a16

Aperta
#1,450 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
45/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
machine-learning, python

Direzione di ricerca

The issue is in winml-cli's reporting logic for compiled models. Start by examining the code that generates the model_info.precision field in the perf command output, likely in a module handling model metadata or report generation. Compare how precision is determined for raw vs. compiled models. Look at the EPContext handling and the Q/DQ initializers. Run the provided reproduction script to see the discrepancy in the generated JSON files.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug P1

Summary

Compiled w8a16 model is reported as w16a16

Priority: P1, confirmed by the release driver on 2026-09-23.
Bugbash reference: 2026-09-22 / BUG-004.
Status: Observed in the first test round; no fix or post-fix regression is recorded. This issue is based on retained test evidence, not a new test run.

Environment

  • Test date: 2026-09-22; findings triaged on 2026-09-23.
  • CLI: winml-cli 0.4.0, installed from a local wheel. The exact source commit/build provenance is not recorded in the test evidence.
  • Hardware: Snapdragon X Elite X1E80100, Qualcomm Hexagon NPU.
  • OS: Windows 11 26H1, build 28000.2956, ARM64.
  • Python: 3.11.15, workspace virtual environment.
  • EP: QNNExecutionProvider from WinML Catalog, package 2.2480.49.0.
  • ONNX Runtime: 1.27.1.202607110137, Windows ML distribution.
  • ONNX: 1.18.0; PyTorch: 2.14.0; Transformers: 4.57.6.
  • Model: catalog microsoft/resnet-50, image-classification, pixel_values FP32 [1, 3, 224, 224].
  • Quantization, where applicable: w8a16, uint8 weights / uint16 activations, 10 calibration samples.

Reproduction

Run in PowerShell with winml-cli 0.4.0 installed. Use a new output directory and retain the external files generated alongside the compiled ONNX.

$results = '.\repro-compiled-precision'
New-Item -ItemType Directory -Path $results | Out-Null
uv run winml export -m microsoft/resnet-50 -o "$results\resnet.onnx"
uv run winml optimize -m "$results\resnet.onnx" --ep qnn --device npu --disable-ort-graph-optimization -o "$results\resnet-opt-no-ort.onnx"
uv run winml quantize -m "$results\resnet-opt-no-ort.onnx" -p w8a16 --samples 10 --task image-classification --model-id microsoft/resnet-50 -o "$results\resnet-w8a16.onnx"
uv run winml compile -m "$results\resnet-w8a16.onnx" --ep qnn --device npu -o "$results\resnet-compiled.onnx"
uv run winml perf -m "$results\resnet-w8a16.onnx" --ep qnn --device npu --iterations 10 --warmup 2 -o "$results\precision-raw.json"
uv run winml perf -m "$results\resnet-compiled.onnx" --ep qnn --device npu --iterations 20 --warmup 3 -o "$results\precision-compiled.json"

Compare model_info.precision in the two reports. The optimize flag in preparation bypasses a separate serialization issue.

Expected result

Preserve the source precision, or report it as unknown when compiled weights are opaque.

Actual result

Both perf commands exit successfully, but their model_info.precision values disagree:

Input Reported precision
Raw uint8-weight / uint16-activation QDQ model w8a16
Compiled EPContext from that same model w16a16

This discrepancy was independently confirmed from the saved JSON reports.

Workaround

Use the verified pre-compile quantization metadata when interpreting results; there is no verified CLI workaround that corrects the compiled report field.

Scope and evidence

The compiled outer ONNX exposes UINT16 activation-side Q/DQ initializers, while weights are opaque in the external EPContext. This confirms a report-label defect, not an actual conversion of weights to 16 bits. Successful inference/exit code does not make the precision field correct. Preserve known source precision or report unknown when weight precision cannot be established.

The following evidence is retained locally by the reporter; these filenames are an index, not uploaded attachments:

  • quantize-w8a16.stdout.log
  • perf-raw.json
  • perf-compiled.json
  • onnx-validation.json
Lingua principale
Python
Stelle
40
Fork
11
Merge medio
1g 2h
PR unite (30g)
59

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di microsoft/winml-cli

Tutte le issue di microsoft/winml-cli

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.