Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

fix(ci): clean up the Activity gate server and score on a quiet host

Aperta
#1,949 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
45/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
github-actions, python, shell

Direzione di ricerca

The issue is in the CI job defined in .github/workflows/ci.yml, specifically the 'WebUI installed artifact' job and the step at line 738. The main scripts are scripts/webui/verify_activity_performance.py and scripts/webui/verify_activity_performance_tests.py. Start by understanding the server launch and cleanup logic in verify_activity_performance.py, especially the start_server function and terminate_owned. The fix involves adding PID/PGID recording, improving cleanup in the finally block, and extending the precondition to check host load. Verify changes by running the existing tests and observing the CI job behavior.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

priority:medium status:ready type:bug

Problem / Background

The WebUI installed artifact job (.github/workflows/ci.yml:551) fails its Activity performance gate step (:738) for two distinct environmental reasons, and one of them poisons the job's own next run. Both were observed on the self-hosted GB10 runner lablup-dgxspark21 on 2026-09-21 and verified directly on the host. They live in the same step, one PR fixes both, and one verification pass covers both, so implement them together.

Fault 1: the job leaks its server, and the leak blocks the next run. The step opens with a fail-closed precondition at :748-751: it counts GPU compute processes and, if the count is non-zero, emits ::error::GPU compute processes are already present before Activity gate and exits 1 before scoring anything. On PR #1946 that precondition fired, and the single GPU process was the job's own server: pid 1768781, mlxcel-server-webui, started from $RUNNER_TEMP/mlxcel-webui-installed/ (the copy made at :582), holding 4828 MiB. Its parent was init (PPID 1) and its working directory read .../_work/_temp/mlxcel-activity-performance-kj_i_r3k (deleted), so the temp directory it was launched from had already been removed. The owning job had completed hours earlier with conclusion failure. Clearing it by hand with a single kill let the next run pass. The precondition itself is correct and is doing exactly its job; the defect is upstream, in the absence of a cleanup that survives the failure path.

Fault 2: the gate scores under host load and reports a degradation inside its own noise. PR #1944 scored median_decode_degradation_percent: 2.798, paired_range_percent: [-1.147, 4.771], status: investigate, while a second CI job compiled on the same host. PR #1946 scored median_decode_degradation_percent: 3.871, baseline_cv_percent: 3.233, paired_range_percent: [-1.139, 7.548], status: investigate. In both cases the reported range crosses zero, so the measurement does not establish a regression in the direction it is failing on, and on #1946 the median degradation is 1.20x the baseline's own coefficient of variation. Neither PR touched any TypeScript, JavaScript or WebUI file: #1946's nine files were two workflows, docs/installation.md, scripts/bench_block_width.sh, two guard scripts, build.rs, cuda_arch.rs and cuda_arch_tests.rs. Both passed on a re-run, and the #1946 re-run passed specifically because it was held until every other check had left pending and gpu_process_count read 0 before starting. That is the empirical evidence for the fix direction: the gate needs the runner to be quiet, not merely the GPU to be free of compute processes, because a concurrent compile on the same machine is enough to move the score.

Current Behavior

How the server survives the step. scripts/webui/verify_activity_performance.py:439 launches the server with start_new_session=True and cwd=h.work, where h.work comes from unique_work_dir (:104-108, prefix mlxcel-activity-performance-), which is exactly the orphan's observed working directory. Because the server is its own session leader, a process-tree kill of the step does not reach it. The only cleanup is the in-process finally at :686-690, which calls terminate_owned (:371), and that cleanup has three ways to leave something alive. It is skipped entirely when the Python process is killed rather than unwound (the step's timeout-minutes: 45 at :739, job cancellation through the workflow concurrency group at :64, or a runner-level abort). It can abandon itself partway: the SIGKILL fallback's proc.wait(timeout=5) at :389 is not wrapped in except subprocess.TimeoutExpired the way :382 is, so a process that does not die within five seconds raises out of the finally and the remaining cleanup never runs. And it never escalates when the parent dies quickly: SIGKILL at :386 is reached only if proc.wait(timeout=8) at :382 times out, so a parent that exits inside eight seconds leaves forced False with nothing further signalled, while process_group_empty(proc.pid, 0.5) at :393 is computed and recorded into the result dict without anything ever acting on a False. The job has no if: always() step other than the four artifact uploads at :774, :781, :788 and :798.

Which of those three produced the 2026-09-21 orphan could not be determined: the step logs for those runs have expired, and the one identifiable failed job (run 35590282982, job 106302926940) ran its remaining steps normally after the gate step failed, so its finally did execute. Treat the three paths as read from the code rather than observed, and close all three rather than betting on one.

How the verdict is computed. webui/scripts/activity-performance.mjs:133 computes the paired degradation (off - on) / off * 100, :136 the population CV of the five off rates, and :138 sets status: degradation > 2 || cv > 5 ? 'investigate' : 'within-target'. Five alternating pairs at roughly 1.6 s each is the whole sample.

Why GitHub concurrency does not already prevent Fault 2. All six GB10 jobs in ci.yml land on the single runner lablup-dgxspark21 and run strictly sequentially; the job timings of recent runs show no overlap between WebUI installed artifact and the five compile and link jobs that share the runner. The interference therefore comes from outside this workflow's serialization: another runner label registered on the same host (pipeline-parallel-ci.yml:216 uses self-hosted plus pp-three-host, release.yml:521 uses GB10), or host-local work. Adding a concurrency: group cannot fix this; the gate has to measure the host itself.

When Fault 1 fires there are no numbers at all. The precondition exits before python3 scripts/webui/verify_activity_performance.py at :755 is ever invoked, so neither webui-activity-performance-summary.json nor -full.json is written, the uploads at :788 and :798 report No files were found with the provided path under if-no-files-found: warn, and the run carries no degradation figure. A reader who sees that failure should not go looking for one.

Proposed Solution

These are fix shapes offered for review, not settled decisions, except where marked as rejected.

  1. Cleanup that survives the failure path. Have start_server record the server's pid and pgid to a stable path under $RUNNER_TEMP (for example $RUNNER_TEMP/webui-activity-server.pid) as soon as the process starts, and add an if: always() step after the gate that reaps by two independent handles, the recorded process group and any surviving process whose executable is $RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-webui or whose working directory is a mlxcel-activity-performance-* directory, reports whether it had to signal, and fails loudly if it cannot. Confirm first whether the server spawns any child process and whether such a child stays in the server's process group; if a child can leave the group, a pgid-only reap reaches nothing and the identity handle is the one that does the work. In the verifier, guard the wait at :389 with except subprocess.TimeoutExpired so the in-process finally can no longer abandon its own remaining cleanup, and make terminate_owned act on the process_group_empty result it already computes at :393 instead of only recording it. src/server/router_webui_playwright_harness_tests.rs:292-321 already implements the intended shape on the Rust side (spawn with process_group(0), SIGTERM, wait for the group to empty, then escalate to SIGKILL) and is the reference to mirror. The graceful path stays primary: scripts/webui/summarize_activity_evidence.py:build_summary still requires server_shutdown.exit_code == 0, forced false and process_group_empty true, so the workflow step is a fallback for abnormal death, not a replacement for orderly shutdown.

  2. Distinguish a leaked orphan from the current run's own server. The precondition should name what it found rather than only counting it. A leaked orphan is cheaply identified by a /proc/<pid>/cwd that reads (deleted), because the run's temp directory is removed when the job ends, and by a parent chain that does not reach Runner.Worker; the current run's own server has a live cwd and a parent chain that does. That distinction was used successfully during this investigation and is the recommended fallback for a process no pid file claims. Do not try to use /proc/<pid>/environ for this: clean_env at :125-127 passes a fixed nine-entry allowlist that drops every GITHUB_* variable, so the server carries no GITHUB_RUN_ID to match against.

  3. Gate on host quiet, not just on a free GPU. Extend the precondition to wait, on a bound of its own like the existing flock -w 600 at :747, until the host is idle enough to measure, then fail closed with a message naming what it was waiting on. Decide quiet from the runnable-process count in /proc/loadavg or from processes in R state, never from a ps %cpu threshold, because %cpu from ps is a lifetime average and reads a freshly started compile as idle. Record the host state observed at gate start (load average, foreign compiler or runner pids) into the full evidence JSON so a noisy verdict can be attributed after the fact instead of re-derived from scratch.

  4. Rejected: loosening the 2 percent median or 5 percent CV threshold. The numbers move because of when they are taken, so a wider threshold would hide real regressions rather than fix the measurement.

Relationship to #1925, read this before implementing. #1925 is open and already claimed, and it owns the verdict rule itself: a confidence interval in place of the bare median, a distinct inconclusive status for an undecidable measurement, baseline CV demoted from failure condition to diagnostic, and one shared fixture exercising the JS and Python sides. That work delivers the three-way "no regression" / "unresolved within the measurement floor" / "regression" reporting and keeps paired_range_percent and baseline_cv_percent in the summary, so do not re-specify or re-implement it here. This issue owns the CI step only. The split matches #1925's own scope, which touches .github/workflows/ci.yml "only if the timeout needs adjusting". Do not modify webui/scripts/activity-performance.mjs or the verdict thresholds inside the two Python validators as part of this issue, or two implementers will collide on the same lines.

Scope

In scope: .github/workflows/ci.yml (the webui-installed-artifact job: the gate step's precondition and a new always-run cleanup step), scripts/webui/verify_activity_performance.py (pid and pgid recording in start_server at :428-446, the unguarded wait at :389, host-state capture into the evidence), and scripts/webui/verify_activity_performance_tests.py.

Out of scope: the verdict rule, statistic and thresholds (#1925); the deferred native hidden acceptance; runner provisioning or a second GB10 runner; anything under webui/src or src/server.

Implementation Notes

  • Reuse: keep the fail-closed precondition at :748-751 and the cooperative flock at :745-747 and extend them rather than replacing them; reuse process_group_empty (:396) and stop_process_group (:408) for pgid liveness instead of writing new equivalents; source scripts/ci/_common.sh if the cleanup shell grows past a few lines.
  • Constraints: the step must stay inside timeout-minutes: 45, and the quiet-host wait needs its own bound so a permanently busy host fails with a clear message instead of consuming the whole budget. The pid file must live under $RUNNER_TEMP, not inside the verifier's per-run mkdtemp work directory, because that directory is precisely what disappears.
  • Edge cases: no pid file, meaning the gate never reached start_server, is a clean no-op rather than a failure; a recorded pid that the kernel has since reused must not be killed, so check the executable name or start time before signalling; a server already gone when cleanup runs reports that it had nothing to do; an orphan the cleanup cannot kill fails the step rather than being left for the next run to trip over.
  • Error handling: the cleanup step reports pid, pgid, whether a signal was needed, and whether the process group was empty afterwards. The precondition names each foreign GPU process (pid, command, whether its cwd reads (deleted), memory held) so a reader can tell a leak from a genuine conflict without logging into the runner.

Acceptance Criteria

  • A deliberately failed gate run leaves no GPU compute process behind: after the job completes, nvidia-smi --query-compute-apps=pid and pgrep -f mlxcel-server-webui both return empty on lablup-dgxspark21, with the run ID recorded in the PR body.
  • The cleanup runs on the failure path and reports what it did, while the passing path still records server_shutdown.exit_code == 0, forced: false and process_group_empty: true, so summarize_activity_evidence.py:build_summary accepts the summary unchanged.
  • A leaked orphan seeded by hand before a run is identified as a leak and named with its pid, command and deleted cwd, and is distinguished from the current run's own server rather than only counted.
  • terminate_owned cannot raise out of its own SIGKILL fallback: a test in scripts/webui/verify_activity_performance_tests.py using a process that ignores SIGTERM and outlives the wait shows the function returning a result with forced: true instead of propagating subprocess.TimeoutExpired. A second test covers the escalation gap: a parent that exits inside the eight second wait while leaving a live process in its group must end with the group empty, which current code does not achieve because :393 records process_group_empty without acting on it. Both tests fail on current code.
  • The gate refuses to score while the host is busy: with a synthetic load running on the runner, the step waits and then fails closed naming host load, and with the load removed the same commit passes.
  • The full evidence JSON records the host state observed at gate start, and the summary still passes scripts/webui/summarize_activity_evidence.py.
  • The gate still runs inside the real WebUI installed artifact job on every WebUI or Rust change, and the change is integrated into that job rather than left as a standalone script.
  • Everything above lands in a single PR, with the verification block run once at the end and pushes kept to a minimum, because every push touching this job costs a full GB10 gate cycle.

Deliberately not an acceptance criterion here: the three-way "no regression" / "unresolved within the measurement floor" / "regression" reporting, with paired_range_percent and baseline_cv_percent retained in the summary, belongs to #1925 and must not be duplicated in this PR.

Verification

python3 scripts/webui/verify_activity_performance_tests.py
python3 scripts/webui/summarize_activity_evidence_tests.py
make verify-webui-helper-tests
# On lablup-dgxspark21, immediately after a deliberately failed gate run:
nvidia-smi --query-compute-apps=pid,used_gpu_memory --format=csv
pgrep -af mlxcel-server-webui

A pass is: helper tests green; no GPU compute process and no mlxcel-server-webui left behind after the deliberately failed run; the quiet-host wait visible in the job log under synthetic load and the same commit green once the load is removed; and a green WebUI installed artifact run on an otherwise unchanged tree.

Technical Considerations

Related: #1925 (verdict rule, open and claimed; coordinate before touching the scoring path). lablup-dgxspark21 serves six ci.yml jobs plus the GB10 job in release.yml:521, so an orphaned server holding 4.8 GiB of GPU memory for hours blocks far more than this one gate, which is why the cleanup should fail loudly when it cannot reap rather than reap silently.

Lingua principale
Rust
Stelle
471
Fork
55
Merge medio
9h 36m
PR unite (30g)
289

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di lablup/mlxcel

Tutte le issue di lablup/mlxcel

Issue simili

Altre issue su Rust

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.