Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

CI flakiness inventory: Sep 19-26, 2026 (follow-up to #12704)

Abierto
#12,929 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
25/100
Tipo de issue
Error
Claridad
Necesita aclaración
Estado de actividad
Activo
Stack tecnológico
android, azure, csharp

Línea de trabajo

Start by reading the unresolved-signature evidence and the linked trackers, especially #12658, #12652, #12916, and #12927. Inspect the cited build logs, dotnet-run-device-output.log, and the publish binlog where applicable; done means separating correlated failures from independent flakes and recording evidence without treating every named test as a distinct issue.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

flaky-ci flaky-tests needs-triage
Android framework version

net11.0-android (Preview); the CI matrix also exercises net10.0-android test flavors.

Affected platform version

dotnet/android public Azure DevOps dotnet-android pipeline (dnceng-public, definition 333), macOS/Linux/Windows. This follow-up covers builds queued 2026-09-19 21:25:43 through 2026-09-26 21:25:43 UTC.

Description

Follow-up to the closed previous flaky-CI inventory #12704. That issue covered PR builds from the earlier interval; this inventory includes PR, scheduled main, and manual builds, so its percentages are not directly comparable with #12704.

Population at the cutoff: 241 builds (232 PR, 7 scheduled, 2 manual): 94 succeeded, 145 failed, 1 canceled, 1 in progress. The in-progress build 1613531 was canceled after the cutoff. The 239 completed, non-canceled builds published 18,191 test runs and 2,754,737 test executions, including 4,387 failed-result entries spanning 550 automated test names. Nine green builds still published 27 non-gating failed results; 65 red builds had no published failed-test result. These are not 550 independently flaky tests: #12903 alone accounts for 2,766 failed-result entries (63.0%), and #12903/#12887/#12889 together account for 3,848 (87.7%), much of it correlated with changes on those branches.

Counts below are distinct affected builds per signature, unless explicitly called results or tasks. One build may have several causes; warning counts, test-result counts, and red-build counts must not be added together. An otherwise-green build with a failed test is a lead, not by itself proof that the case passed on retry. Source builds and the prior issue are linked for reproducibility.

Ongoing or resurfaced since #12704
Signature This interval Prior attempt / current assessment
Helix device install/UID 9 affected builds: 8 published failed WorkItemExecution results, plus a UID failure in 1607534 with no published failed work item. One of the eight, green 1610810, passed on a later Helix attempt. The device-lab problem from #12704 remains under OPEN #12658 (and dotnet/arcade#17483). Recovery proposal OPEN #12666 uses XHarness; CLOSED unmerged #12715 attempted another recovery. MERGED #12641 fixed a different, fork-authentication 401 path. Do not group the four app crashes below with UID failures.
Hosted-agent disconnect / worker timeout 2 lost-heartbeat builds (1606715, 1610971) and 2 worker-timeout builds (1611229, 1612872). Agent disconnections appeared in #12704. No corrective repository PR found; inspect the agent/pool and rerun only affected jobs. The underlying cause of each timeout remains unknown.
Azure download/feed DNS, Git checkout transport, cache 429 2 artifact DNS failures (1609832, 1613231), 1 tool-feed DNS failure (1605734), 2 GitHub submodule transport failures (1606121, 1608238), and 1 cache HTTP 429 (1607186). Artifact connectivity, a submodule failure, and cache rate limiting were reported in #12704. No DNS/transport root fix found. CLOSED unmerged #12705 proposed bounded cache-task retries; checkout had already retried on one failure. Task-level retry is a mitigation, not a network fix.
JNI global-reference leak assertion TryFindClass_Utf8_DoesNotLeakGlobalRefs failed in 2 builds, with a later same-flavor pass in green 1611330. #12704 recorded the String overload and other process-wide GREF assertions. MERGED #12032 disabled UTF-8 temporarily; CLOSED unmerged #12038 proposed a dedicated harness; MERGED #12709 isolated/re-enabled the test, yet this week's off-by-one failures followed it. Tracker OPEN #12031 remains relevant. Do not silently relax the assertion and mask real leaks.
Gradle Facebook device test GradleFBProj failed non-gating in two unrelated green builds, 1605257 and 1608189. Three more failures were on code-changing #12903; their generic FailedBuildException is not independently classified as a flake. Same test family appeared in #12704. CLOSED unmerged #12719 proposed bounded retries only for diagnosed transient dependency/SDK errors; MERGED #12198/#12199/#12707/#12902 address mirrors, Gradle state, or wrapper downloads, but do not establish a fix for both current cases. Inspect attached build logs before reusing #12719's classifier.
Device-run output DotNetRunWithDeviceParameter: 4 failed entries / 4 builds, including 2 green builds (1607515, 1609872). Earlier tracker #10832 CLOSED when failures went quiet; #11320 also recorded it. MERGED #10740 originally wired --adb-target, but the missing-serial/message assertions resurfaced. Inspect dotnet-run-device-output.log before distinguishing product argument loss from emulator output timing.
New to this inventory, or newly specific and still unresolved
Signature This interval Status / next evidence
macOS PowerShell host startup, procargs errno 5 6 builds at the original cutoff (1607771, 1608663, 1609382, 1610691, 1613229, 1613531); five red and one then active. The sixth error was captured at 21:23 UTC before 1613531 was later canceled and its log became unavailable. Already tracked separately in OPEN #12652, but absent from #12704's inventory. Different tasks fail before script execution; no corrective PR found. Investigate macOS agent/pwsh startup, not the individual script; targeted retry is only a mitigation.
Android SDK android-36.1 removal, MSB3231 4 red builds: 1607212, 1610249, 1612828, 1612840; data/res remained in the directory despite three setup attempts each. CLOSED unmerged #12767 proposed staged/atomic replacement but had unresolved backup-cleanup and validation review findings. Rework that design; simply increasing retries did not work.
Hard test timeouts without useful failed results 25 timed-out test tasks in 6 builds, including the same nine Windows tasks in each of 1612401 and 1612859. Separately, 10 jobs in 4 builds hit a 180-minute job cap (7 in 1605352). #12916 and dependent #12927 are both OPEN and contain, not fix, the 18 Windows timeouts. The same shards stopped progressing early in both builds. Diagnose their shared testhost/build-process wait; do not presume 25 independent infrastructure failures or raise limits blindly. Other job-cap causes remain unclassified.
Individual test candidates PingLocalhost timed out then passed the same flavor in green 1607347; BuildBasicApplicationThenMoveIt(True,CoreCLR) had 4 file-sharing violations in 4 distinct PR builds; NativeAOTSample had 5 failures / 5 builds, one green; CoreCLRAssemblyNameWithNativeImageSuffix failed then passed amid 18 failures in one shared retry lane of 1607515. Ping and the JNI result above have direct intermittent evidence; the file-lock test is a likely contention flake. NativeAOTSample needs logcat before attributing an app-start failure to emulator timing. Treat the 18-result retry lane as one correlated incident, not 17 independent flaky tests. No direct fix PR was found for these new candidates.
Azure Test/telemetry publication degradation Test REST run-summary: 70 builds / 692 retry warnings; telemetry: 45 builds / 142 DNS or timeout warnings; actual test-result publication warnings: 3 builds. Newly counted separately from gating failures. Twenty telemetry-warning builds were green. Determine whether results were actually absent before filing these warning-only symptoms as Azure service outages; no repository fix PR found.
Scheduled Ref.37 pack assertion XASdkTests.DotNetPublish: 12 failing cases in 3 scheduled main builds (1604684, 1608359, 1611817). Six matching cases passed an earlier build of the same SHA. Not established as missing-pack infrastructure or an individual flaky test. The assertion enumerates an existing Microsoft.Android.Ref.37/.../Mono.Android.dll path but fails to match the expected publish output. MERGED #12359 introduced this assertion; compare its match logic and publish binlog before changing pack installation.
Possible recurrence of a previously fixed peer-identity bug GetObjectArray failed in 1605734 with “Expected existing Context peer, got ... App,” after MERGED #12650. #12704 had classified #12650 as likely fixed. This build is on #12851, which changes TrimmableTypeMap, so determine PR-related regression versus persistent flake before reclassifying the fix. The older tracker #10973 is still open.
Fixed or mitigated signatures: distinguish the exact cause from the test name
Prior/current intervention Evidence and status
MERGED #12886, corrupt cached SDK archive recovery The 1608359 android_m2repository_r47.zip bad-CRC extraction preceded the merge on Sep 23. No further verified same-archive extraction failure was found later in this window. Recovery applies when the cached ZIP's SHA differs from its pinned SHA; do not claim every unzip error is solved.
MERGED #12641, Helix fork-token authentication The earlier fork HTTP 401 path was fixed; this week's UID-assignment failures and #12847 app crashes are different Helix failures. The broader Helix issue remains ongoing.
Previously classified likely fixed in #12704: #12515 apkdiff timeout, #12649 missing SharpZipLib/Polly, #12603 stale JAVA_HOME, #12700 Roslyn MemorySafetyRulesAttribute Some same-named tests failed this week for different observed reasons: e.g. 27 builds reported BuildReleaseArm64 failures involving missing linker outputs, not the old apkdiff timeout; InstallAndroidDependenciesTest reported missing SDK source.properties, not missing SharpZipLib/Polly; CheckSignApk reported build failures without the old stale-JAVA_HOME signature. Do not count these as proof that the earlier specific fixes regressed, or as proof of a new flake without their logs.
MERGED #12903, after its earlier red PR builds Its later PR build 1612266 succeeded. The 2,766 earlier failed entries were primarily correlated with a changing workload-dependencies PR and are excluded from flaky-test totals. Compatibility work remains in OPEN #12923.

Unclassified/excluded rather than declared fixed: FastTimingTests.ConcurrentEventsCanGrowAndDump still appears on code-changing PR builds after MERGED #12492; its current error may differ from the old logcat race and needs logs. The 21 generated-Java cannot find symbol build failures on the still-open #12890–#12895 stack are product/stack failures, not infrastructure flakes. A previous BuildArtifactsOutputPaths test fix #11839 CLOSED unmerged; this week's cases fail during build, so that earlier missing-output-path fix cannot be assumed sufficient. OPEN #12926 improves classification and targeted retries, but does not itself fix any root cause above.

Proposed follow-up
  • Re-evaluate #12650/#10973 against build 1605734 on #12851; preserve the distinction between PR regression and recurring flake.
  • Rework #12767's atomic SDK replacement with the review issues addressed; follow #12652 for the pwsh host failures.
  • Continue #12658's device-lab work and evaluate OPEN #12666 XHarness recovery. The four other Helix work-item failures, all in build 1605352, are app crashes from missing jni_remapping_type_replacement_count; their correction is already in OPEN #12847 and must not be charged to device-lab UID failures.
  • Investigate the #12916/#12927 Windows testhost hang and the Ref.37 assertion using task logs/binlogs before changing timeout limits or workload installation.
  • Revisit #12031's leak-test isolation after the two post-#12709 failures; update #10832 with the new DotNetRunWithDeviceParameter output logs. Consider only classified targeted retries for confirmed external errors.
Steps to Reproduce
  1. Enumerate public dotnet-android definition 333 builds queued in the exact UTC window above, recording each build's status at the cutoff. Include scheduled main and manual runs as well as PRs; canceled build 1613531 may no longer appear in a later default build-list query.
  2. For each build, read its Azure timeline, failed/canceled job and task issues, and the root task log—not just the GitHub aggregate check or failed-test API.
  3. Query ResultsByBuild?outcomes=Failed, the build's test runs, and the failed run's exact result title/error. Compare repeated results in the same configuration and unchanged-source reruns; green results can contain non-gating failed attempts.
  4. Compare the exact error signature against #12704 and the linked fix PR, not merely the test method name. Check changed files and stacked PR bases before attributing correlated test failures to infrastructure.
  5. Reproduce specific examples using the linked Azure builds. For device crashes or hangs, inspect the matching published logcat/Helix work-item artifacts; for generic FailedBuildException, inspect the attached build log/binlog before selecting a retry or a code fix.
Did you find any workaround?

A targeted rerun of the affected job/stage can unblock a confirmed transient DNS, hosted-agent, or device-lab failure, but it must not treat deterministic PR regressions, all Helix work-item failures, or all generic test failures as flaky. #12926 is an OPEN proposal for classified, confirmation-gated targeted retries, not a landed root-cause fix. #12705's CLOSED unmerged bounded cache-task retry is a small mitigation for HTTP 429. For SDK cleanup, three task attempts already failed: the unmerged atomic-replacement design in #12767 needs correction rather than more retries. For Helix UID failures, #12666 proposes XHarness recovery while device-lab cleanup is investigated.

Relevant log output
Call to 'procargs' failed with errno 5
error MSB3231: Unable to remove directory ".../platforms/android-36.1". Directory not empty: '.../data/res'
DownloadPipelineArtifact: nodename nor servname provided, or not known (dev.azure.com:443)
GetObjectArray: Expected existing Context peer, got Android.AppTests.App

Full task logs, case-level errors, and retries are linked to representative builds in the Description; the lines above are only short identifying signatures.

Lenguaje dominante
C#
Estrellas
2.1k
Forks
580
Merge medio
1 d 22 h
PR fusionados (30 d)
215

Preparar el entorno

Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de dotnet/android

Todos los issues de dotnet/android

Issues similares

Más issues de C#

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.