[winml perf] QNN/NPU basic op-tracing produces no profiling CSV
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 45/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- cli, machine-learning, python
- Área
- cli, machine-learning, performance, tooling
Línea de trabajo
The issue is in winml-cli's perf command with --op-tracing basic for QNN/NPU. Start by examining the CLI's tracing implementation, likely in a command/perf.py or similar module. Look for where the CSV is generated and why it might be empty. Check the integration with ONNX Runtime's QNNExecutionProvider and ETW sessions. Run the provided reproduction steps to see the error code 4. 'Done' is when the command produces a profiling CSV for the given models.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
QNN basic operator tracing produces no profiling CSV
Priority: P1, confirmed by the release driver on 2026-09-23.
Bugbash reference: 2026-09-22 / BUG-002.
Status: Observed in the first test round; no fix or post-fix regression is recorded. This issue is based on retained test evidence, not a new test run.
Environment
- Test date: 2026-09-22; findings triaged on 2026-09-23.
- CLI: winml-cli 0.4.0, installed from a local wheel. The exact source commit/build provenance is not recorded in the test evidence.
- Hardware: Snapdragon X Elite X1E80100, Qualcomm Hexagon NPU.
- OS: Windows 11 26H1, build 28000.2956, ARM64.
- Python: 3.11.15, workspace virtual environment.
- EP: QNNExecutionProvider from WinML Catalog, package 2.2480.49.0.
- ONNX Runtime: 1.27.1.202607110137, Windows ML distribution.
- ONNX: 1.18.0; PyTorch: 2.14.0; Transformers: 4.57.6.
- Model: catalog microsoft/resnet-50, image-classification, pixel_values FP32 [1, 3, 224, 224].
- Quantization, where applicable: w8a16, uint8 weights / uint16 activations, 10 calibration samples.
Reproduction
Run in PowerShell with winml-cli 0.4.0 installed. Use a new output directory. Prepare raw w8a16 and compiled inputs with the same artifact chain used in the bugbash:
$results = '.\repro-qnn-optrace'
New-Item -ItemType Directory -Path $results | Out-Null
uv run winml export -m microsoft/resnet-50 -o "$results\resnet.onnx"
uv run winml optimize -m "$results\resnet.onnx" --ep qnn --device npu --disable-ort-graph-optimization -o "$results\resnet-opt-no-ort.onnx"
uv run winml quantize -m "$results\resnet-opt-no-ort.onnx" -p w8a16 --samples 10 --task image-classification --model-id microsoft/resnet-50 -o "$results\resnet-w8a16.onnx"
uv run winml compile -m "$results\resnet-w8a16.onnx" --ep qnn --device npu -o "$results\resnet-compiled.onnx"
Reproduce on the raw quantized model:
uv run winml perf -m "$results\resnet-w8a16.onnx" --ep qnn --device npu --iterations 3 --warmup 1 --op-tracing basic -o "$results\optrace-raw.json"
Reproduce on the compiled model, including with memory monitoring disabled:
uv run winml perf -m "$results\resnet-compiled.onnx" --ep qnn --device npu --iterations 10 --warmup 2 --no-memory --op-tracing basic -o "$results\optrace-compiled.json"
The disabled ORT graph optimization in preparation bypasses a separately recorded standalone optimize failure; it is not a tracing workaround.
Expected result
Produce basic per-operator profiling data for the QNN/NPU benchmark.
Actual result
Both tracing commands exit with code 4 and produce no profiling CSV:
Op-tracing produced no profiling data (no CSV written).
Ordinary perf without --op-tracing basic has passing records for both raw and compiled inputs.
Workaround
No verified workaround that produces operator-tracing data in this bugbash. Ordinary perf runs successfully, but is not equivalent to operator-level profiling.
Scope and evidence
The failure is observed on this host. Root-cause ownership between CLI, WinML/ORT, and QNN EP remains unconfirmed. The CPU-fallback suggestion in the error is not evidence that CPU fallback occurred. ETW session prerequisites/state were not established by this report; do not infer them from the missing CSV.
The following evidence is retained locally by the reporter; these filenames are an index, not uploaded attachments:
perf-qnn-optrace.stdout.logperf-qnn-optrace.stderr.logperf-compiled-optrace.stdout.logperf-compiled-optrace.stderr.log
- Lenguaje dominante
- Python
- Estrellas
- 40
- Forks
- 11
- Merge medio
- 19 h 32 min
- PR fusionados (30 d)
- 51
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/winml-cli
-
bug P1
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
bug P1 triaged
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
microsoft/winml-cli#1097 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
P2 refactor triaged
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
bug P1
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
Los mantenedores suelen responder en 1 día
-
bug P1
Dificultad 3/5 1-2 días Aptitud para principiantes 65/100
Los mantenedores suelen responder en 1 día
Todos los issues de microsoft/winml-cli
Issues similares
-
bug status/needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
prowler-cloud/prowler#12887 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
use-agent-os/agent-os#3484 ·
Los mantenedores suelen responder en 2 días
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
open-telemetry/opentelemetry-python-contrib#5113 · 2 comentarios · 2 reacciones ·
Los mantenedores suelen responder en 1 día
-
external
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
langchain-ai/docs#6255 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día