ci: the azure-mshv-scus runners fail the MSHV performance gate on guest-exit teardown and 512 MiB restore
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 38/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- azure, linux, rust
- Domain
- ci-cd, performance
Research direction
Read doc/benchmarks.md first, especially its description of separate benchmark series, and compare the MSHV runner labels and host-kernel details in the issue. Decide whether to isolate the scus runners or align their environment with the baseline runners; confirm the chosen approach with repeatable measurements. Done means runner placement no longer causes false performance-gate failures or contaminates the baseline.
Written by the indexing model from the issue text.
Description
Summary
The performance gate fails whenever the Platform / Linux / MSHV / Virtual machine job runs on azure-mshv-scus-1 or azure-mshv-scus-z1-1, whatever the pull request changes. Those runners measure MSHV guest-exit teardown about 15 to 29 ms slower, and the 512 MiB shell snapshot restore at 4 and 8 vCPUs about 20 to 40 ms slower, than the runners that measured the baseline window (azure-mshv-1, -5, -6, and -7). Every rerun that lands on another MSHV runner passes the gate.
#394 and #404 reported the same pattern: those runners use another scale set and host kernel (6.6.137.mshv2-2.azl3, against 6.6.148-200.azl3). Both runners also run Rust 1.99.0, so OpenVMM vmm-tests / Linux / MSHV fails there too (#402).
Evidence
The MSHV platform job's runner and the gate's result in recent pull-request runs:
| Run | Pull request | MSHV platform runner | Gate |
|---|---|---|---|
| 37410478636 | #401 | azure-mshv-scus-1 |
failure |
| 37419824754 | #401 | azure-mshv-7 |
success |
| 37409832176 | #394 | azure-mshv-1 |
success |
| 37439708257 | #404 | azure-mshv-5 |
success |
| 37448382849 | #404 | azure-mshv-1 |
success |
| 37518293095, attempts 1 to 3 | #411 | azure-mshv-scus-z1-1, then azure-mshv-scus-1 twice |
failure |
| 37518293095, attempt 4 | #411 | azure-mshv-6 |
success |
| 37520322807, attempts 1 and 2 | #412 | azure-mshv-scus-1 |
failure |
The same metrics fail in each of the five failing attempts of #411 and #412:
openvmm_cold_start_guest_exit_teardown: 40 to 48 ms, against a base median of 24.94 ms.openvmm_snapshot_restore_guest_exit_teardown: 30 to 37 ms, against 8.16 ms.shell_snapshot_restore_512_mib: 69 to 71 ms at 4 vCPUs (base 46.25 ms), and 91 to 95 ms at 8 vCPUs (base 52.95 ms).
Effect
Any MSHV runtime pull request fails the gate when its platform job lands on these runners. CI routes by labels, which all MSHV runners share, so the only remedy is to rerun until the job lands elsewhere. The gate cannot tell a regression from runner placement. dev pushes that ran there, such as 37503416598 for #404's merge, may also feed the baseline history with these slower samples.
Options
- Give the
scusrunners a series of their own, asdoc/benchmarks.mddescribes for AMD runners, or keep the platform jobs off them, for example with an extra runner label for the gate's platform jobs. - Or align their host kernel and scale set with the baseline runners, and confirm that their measurements then match.
[!NOTE]
Runner naming update (2026-10-07): Names in linked historical jobs and dated tables remain exactly as GitHub reported them. Current sequential aliases and runner registrations are mapped by VM instance in #211's rename note. Use the current name when connecting to or scheduling a runner.
- Dominant language
- Python
- Stars
- 237
- Forks
- 14
- Avg merge
- 13h 18m
- Merged PRs (30d)
- 232
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/nvx
-
Managed start can fail when OpenVMM reads its control capability before NVX writes itPossibly taken @ppenna claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
confirmed
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day
-
Fall back from --cpu-profile auto to a host profile, with a warning, on CPUs that no built-in profile servesPossibly taken @ppenna claimed this 1 day ago. Open
microsoft/nvx#434 · 1 assignee ·
Maintainers usually reply within 1 day
-
CI: maximize workflow parallelism and reduce turnaround timePossibly taken @ppenna claimed this 1 day ago. Open
Difficulty 5/5 Over a week Newbie friendliness 35/100
microsoft/nvx#430 · 1 assignee ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
microsoft/nvx#426 · 1 comment ·
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
NVIDIA/earth2studio#1241 ·
Maintainers usually reply within 3 days
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Opendocumentation priority: low size: S
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Maintainers usually reply within 1 day
-
bug help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 65/100
ansys/pydpf-core#3547 ·
Maintainers usually reply within 1 day
-
good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
OktoLabsAI/okto-pulse#114 ·
Maintainers usually reply within 1 day