BUG SDK JUnit results are never published because the runner OS condition is wrong
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
Research direction
Start with .github/workflows/build_and_test.yml, job main-job, and the Publish Pytest Results step; compare the matrix OS labels with the runner.os values and review the existing pinned publisher action. Validate success, test-failure, cancellation, and missing-XML paths, ensuring raw JUnit artifacts remain downloadable, test failures stay authoritative, and fork/security boundaries are preserved.
Written by the indexing model from the issue text.
Description
Describe the bug
The Publish Pytest Results step in .github/workflows/build_and_test.yml uses:
if: runner.os == 'ubuntu-latest'
ubuntu-latest is a runner label from matrix.os; the corresponding runner.os value is Linux. The condition therefore never matches. In the audited September 21-22 main-branch window, all 504 SDK result-publisher steps were skipped, across successful, failed, and cancelled matrix jobs. This does not mean the tests themselves were all skipped.
The workflow also has no independent upload of raw SDK JUnit XML. This makes real failures harder to inspect, including the MOSSBench setup error in job 106560764233, whose workflow aggregate is cancelled even though the test job failed.
A second boundary matters when correcting the OS comparison: a plain step condition without an explicit status check normally inherits success-only gating. We need reports from failed test steps too, not just successful runs.
Steps/Code to Reproduce
- Inspect
.github/workflows/build_and_test.yml, jobmain-job, stepPublish Pytest Results. - Compare the
matrix.osvalues (ubuntu-latest,windows-latest,macos-latest) with therunner.oscontext (Linux,Windows,macOS). - Inspect main run 35668881112, attempt 1. The unit-test job writes
junit/test-results.xml, but its publisher step is skipped and no independent SDK JUnit artifact is exposed.
The condition itself is a deterministic reproduction; no live target, new CI run, or probabilistic test failure is needed to establish the mismatch.
Expected Results
Publish SDK test results for the intended supported runners whenever test output exists, including after unit-test failure, and retain downloadable raw JUnit evidence.
Acceptance criteria:
- Compare the correct context/value while preserving the intended supported-runner policy, for example Linux publication using
runner.os == 'Linux'rather than a runner-label string. - Add explicit outcome gating so failed tests do not suppress their result publication. Handle cancellation or missing XML honestly rather than manufacturing a successful test result.
- Retain raw JUnit XML through an independent artifact-upload path, including failed jobs where output exists, with matrix-safe artifact names.
- Keep the unit-test exit status authoritative. Reporting must not hide the failure or make the test step continue on error.
- Preserve fork/security boundaries and least-privilege permissions; do not solve publication by executing untrusted PR code in a privileged event.
- Validate the success, test-failure, and missing-output paths without changing test selection, assertions, thresholds, or retry policy.
This is a workflow reporting repair, not a fix for the underlying MOSSBench or Azure E2E failures. Use existing pinned actions and repository conventions.
Actual Results
Fixed audit window: September 21, 2026, 13:11:11 UTC through September 22, 2026, 13:11:11 UTC.
- 21
build_and_testmain runs; 504 SDK matrix jobs. - Matrix-job outcomes: 422 success, 1 failure, 81 cancelled.
Publish Pytest Results: skipped in all 504 jobs. Some cancelled jobs never reach it, but the OS condition also prevents it in completed Linux jobs.- No independent SDK JUnit artifact upload in the workflow.
- The failed MOSSBench job mentions generated
junit/test-results.xmlin its log, but the report is not available as a standalone artifact.
Screenshots
N/A. The workflow source and linked run/job steps provide the evidence.
Versions
- Workflow:
.github/workflows/build_and_test.yml,main-job. - Matrix: Ubuntu/Windows/macOS, Python 3.11-3.14, default and all-extras installations.
- Observed main snapshot:
2016c4a8566bd66253d431ff38400bade4c77fa3; the same condition remains on main at filing. - Publisher:
EnricoMi/publish-unit-test-result-action, pinned tod0a4676d0e0b938bc201470d88276b7c74c712b3(v2.24.0). - Browser and local package-version snapshot: N/A for this GitHub Actions expression defect.
Related observed failure: #2775 tracks the MOSSBench setup error. This issue repairs reporting and must not hide or change that test failure.
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 896
- Avg merge
- 3d 8h
- Merged PRs (30d)
- 191
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/PyRIT
-
Difficulty 3/5 1-2 days Newbie friendliness 65/100
-
Bug: triage GUI help wanted
Difficulty 3/5 1-2 days Newbie friendliness 65/100
-
feature-request
Difficulty 3/5 1-2 days Newbie friendliness 70/100
-
not ready yet
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
not ready yet
Difficulty 4/5 3-5 days Newbie friendliness 35/100
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
xinnan-tech/xiaozhi-fde-talk#263 ·
-
rules
Difficulty 1/5 Under an hour Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
huggingface/Repo2RLEnv#163 · 1 comment ·
-
Difficulty 1/5 Under an hour Newbie friendliness 95/100
huggingface/sentence-transformers#4074 ·
-
comp/dashboard invalid P3
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
NousResearch/hermes-agent#121143 ·