fix(ci): clean up the Activity gate server and score on a quiet host
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 45/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Ambito
- ci-cd, performance, testing-qa
Direzione di ricerca
The issue is in the CI job defined in .github/workflows/ci.yml, specifically the 'WebUI installed artifact' job and the step at line 738. The main scripts are scripts/webui/verify_activity_performance.py and scripts/webui/verify_activity_performance_tests.py. Start by understanding the server launch and cleanup logic in verify_activity_performance.py, especially the start_server function and terminate_owned. The fix involves adding PID/PGID recording, improving cleanup in the finally block, and extending the precondition to check host load. Verify changes by running the existing tests and observing the CI job behavior.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem / Background
The WebUI installed artifact job (.github/workflows/ci.yml:551) fails its Activity performance gate step (:738) for two distinct environmental reasons, and one of them poisons the job's own next run. Both were observed on the self-hosted GB10 runner lablup-dgxspark21 on 2026-09-21 and verified directly on the host. They live in the same step, one PR fixes both, and one verification pass covers both, so implement them together.
Fault 1: the job leaks its server, and the leak blocks the next run. The step opens with a fail-closed precondition at :748-751: it counts GPU compute processes and, if the count is non-zero, emits ::error::GPU compute processes are already present before Activity gate and exits 1 before scoring anything. On PR #1946 that precondition fired, and the single GPU process was the job's own server: pid 1768781, mlxcel-server-webui, started from $RUNNER_TEMP/mlxcel-webui-installed/ (the copy made at :582), holding 4828 MiB. Its parent was init (PPID 1) and its working directory read .../_work/_temp/mlxcel-activity-performance-kj_i_r3k (deleted), so the temp directory it was launched from had already been removed. The owning job had completed hours earlier with conclusion failure. Clearing it by hand with a single kill let the next run pass. The precondition itself is correct and is doing exactly its job; the defect is upstream, in the absence of a cleanup that survives the failure path.
Fault 2: the gate scores under host load and reports a degradation inside its own noise. PR #1944 scored median_decode_degradation_percent: 2.798, paired_range_percent: [-1.147, 4.771], status: investigate, while a second CI job compiled on the same host. PR #1946 scored median_decode_degradation_percent: 3.871, baseline_cv_percent: 3.233, paired_range_percent: [-1.139, 7.548], status: investigate. In both cases the reported range crosses zero, so the measurement does not establish a regression in the direction it is failing on, and on #1946 the median degradation is 1.20x the baseline's own coefficient of variation. Neither PR touched any TypeScript, JavaScript or WebUI file: #1946's nine files were two workflows, docs/installation.md, scripts/bench_block_width.sh, two guard scripts, build.rs, cuda_arch.rs and cuda_arch_tests.rs. Both passed on a re-run, and the #1946 re-run passed specifically because it was held until every other check had left pending and gpu_process_count read 0 before starting. That is the empirical evidence for the fix direction: the gate needs the runner to be quiet, not merely the GPU to be free of compute processes, because a concurrent compile on the same machine is enough to move the score.
Current Behavior
How the server survives the step. scripts/webui/verify_activity_performance.py:439 launches the server with start_new_session=True and cwd=h.work, where h.work comes from unique_work_dir (:104-108, prefix mlxcel-activity-performance-), which is exactly the orphan's observed working directory. Because the server is its own session leader, a process-tree kill of the step does not reach it. The only cleanup is the in-process finally at :686-690, which calls terminate_owned (:371), and that cleanup has three ways to leave something alive. It is skipped entirely when the Python process is killed rather than unwound (the step's timeout-minutes: 45 at :739, job cancellation through the workflow concurrency group at :64, or a runner-level abort). It can abandon itself partway: the SIGKILL fallback's proc.wait(timeout=5) at :389 is not wrapped in except subprocess.TimeoutExpired the way :382 is, so a process that does not die within five seconds raises out of the finally and the remaining cleanup never runs. And it never escalates when the parent dies quickly: SIGKILL at :386 is reached only if proc.wait(timeout=8) at :382 times out, so a parent that exits inside eight seconds leaves forced False with nothing further signalled, while process_group_empty(proc.pid, 0.5) at :393 is computed and recorded into the result dict without anything ever acting on a False. The job has no if: always() step other than the four artifact uploads at :774, :781, :788 and :798.
Which of those three produced the 2026-09-21 orphan could not be determined: the step logs for those runs have expired, and the one identifiable failed job (run 35590282982, job 106302926940) ran its remaining steps normally after the gate step failed, so its finally did execute. Treat the three paths as read from the code rather than observed, and close all three rather than betting on one.
How the verdict is computed. webui/scripts/activity-performance.mjs:133 computes the paired degradation (off - on) / off * 100, :136 the population CV of the five off rates, and :138 sets status: degradation > 2 || cv > 5 ? 'investigate' : 'within-target'. Five alternating pairs at roughly 1.6 s each is the whole sample.
Why GitHub concurrency does not already prevent Fault 2. All six GB10 jobs in ci.yml land on the single runner lablup-dgxspark21 and run strictly sequentially; the job timings of recent runs show no overlap between WebUI installed artifact and the five compile and link jobs that share the runner. The interference therefore comes from outside this workflow's serialization: another runner label registered on the same host (pipeline-parallel-ci.yml:216 uses self-hosted plus pp-three-host, release.yml:521 uses GB10), or host-local work. Adding a concurrency: group cannot fix this; the gate has to measure the host itself.
When Fault 1 fires there are no numbers at all. The precondition exits before python3 scripts/webui/verify_activity_performance.py at :755 is ever invoked, so neither webui-activity-performance-summary.json nor -full.json is written, the uploads at :788 and :798 report No files were found with the provided path under if-no-files-found: warn, and the run carries no degradation figure. A reader who sees that failure should not go looking for one.
Proposed Solution
These are fix shapes offered for review, not settled decisions, except where marked as rejected.
-
Cleanup that survives the failure path. Have
start_serverrecord the server's pid and pgid to a stable path under$RUNNER_TEMP(for example$RUNNER_TEMP/webui-activity-server.pid) as soon as the process starts, and add anif: always()step after the gate that reaps by two independent handles, the recorded process group and any surviving process whose executable is$RUNNER_TEMP/mlxcel-webui-installed/mlxcel-server-webuior whose working directory is amlxcel-activity-performance-*directory, reports whether it had to signal, and fails loudly if it cannot. Confirm first whether the server spawns any child process and whether such a child stays in the server's process group; if a child can leave the group, a pgid-only reap reaches nothing and the identity handle is the one that does the work. In the verifier, guard the wait at:389withexcept subprocess.TimeoutExpiredso the in-processfinallycan no longer abandon its own remaining cleanup, and maketerminate_ownedact on theprocess_group_emptyresult it already computes at:393instead of only recording it.src/server/router_webui_playwright_harness_tests.rs:292-321already implements the intended shape on the Rust side (spawn withprocess_group(0), SIGTERM, wait for the group to empty, then escalate to SIGKILL) and is the reference to mirror. The graceful path stays primary:scripts/webui/summarize_activity_evidence.py:build_summarystill requiresserver_shutdown.exit_code == 0,forcedfalse andprocess_group_emptytrue, so the workflow step is a fallback for abnormal death, not a replacement for orderly shutdown. -
Distinguish a leaked orphan from the current run's own server. The precondition should name what it found rather than only counting it. A leaked orphan is cheaply identified by a
/proc/<pid>/cwdthat reads(deleted), because the run's temp directory is removed when the job ends, and by a parent chain that does not reachRunner.Worker; the current run's own server has a live cwd and a parent chain that does. That distinction was used successfully during this investigation and is the recommended fallback for a process no pid file claims. Do not try to use/proc/<pid>/environfor this:clean_envat:125-127passes a fixed nine-entry allowlist that drops everyGITHUB_*variable, so the server carries noGITHUB_RUN_IDto match against. -
Gate on host quiet, not just on a free GPU. Extend the precondition to wait, on a bound of its own like the existing
flock -w 600at:747, until the host is idle enough to measure, then fail closed with a message naming what it was waiting on. Decide quiet from the runnable-process count in/proc/loadavgor from processes inRstate, never from aps %cputhreshold, because%cpufrompsis a lifetime average and reads a freshly started compile as idle. Record the host state observed at gate start (load average, foreign compiler or runner pids) into the full evidence JSON so a noisy verdict can be attributed after the fact instead of re-derived from scratch. -
Rejected: loosening the 2 percent median or 5 percent CV threshold. The numbers move because of when they are taken, so a wider threshold would hide real regressions rather than fix the measurement.
Relationship to #1925, read this before implementing. #1925 is open and already claimed, and it owns the verdict rule itself: a confidence interval in place of the bare median, a distinct inconclusive status for an undecidable measurement, baseline CV demoted from failure condition to diagnostic, and one shared fixture exercising the JS and Python sides. That work delivers the three-way "no regression" / "unresolved within the measurement floor" / "regression" reporting and keeps paired_range_percent and baseline_cv_percent in the summary, so do not re-specify or re-implement it here. This issue owns the CI step only. The split matches #1925's own scope, which touches .github/workflows/ci.yml "only if the timeout needs adjusting". Do not modify webui/scripts/activity-performance.mjs or the verdict thresholds inside the two Python validators as part of this issue, or two implementers will collide on the same lines.
Scope
In scope: .github/workflows/ci.yml (the webui-installed-artifact job: the gate step's precondition and a new always-run cleanup step), scripts/webui/verify_activity_performance.py (pid and pgid recording in start_server at :428-446, the unguarded wait at :389, host-state capture into the evidence), and scripts/webui/verify_activity_performance_tests.py.
Out of scope: the verdict rule, statistic and thresholds (#1925); the deferred native hidden acceptance; runner provisioning or a second GB10 runner; anything under webui/src or src/server.
Implementation Notes
- Reuse: keep the fail-closed precondition at
:748-751and the cooperativeflockat:745-747and extend them rather than replacing them; reuseprocess_group_empty(:396) andstop_process_group(:408) for pgid liveness instead of writing new equivalents; sourcescripts/ci/_common.shif the cleanup shell grows past a few lines. - Constraints: the step must stay inside
timeout-minutes: 45, and the quiet-host wait needs its own bound so a permanently busy host fails with a clear message instead of consuming the whole budget. The pid file must live under$RUNNER_TEMP, not inside the verifier's per-runmkdtempwork directory, because that directory is precisely what disappears. - Edge cases: no pid file, meaning the gate never reached
start_server, is a clean no-op rather than a failure; a recorded pid that the kernel has since reused must not be killed, so check the executable name or start time before signalling; a server already gone when cleanup runs reports that it had nothing to do; an orphan the cleanup cannot kill fails the step rather than being left for the next run to trip over. - Error handling: the cleanup step reports pid, pgid, whether a signal was needed, and whether the process group was empty afterwards. The precondition names each foreign GPU process (pid, command, whether its cwd reads
(deleted), memory held) so a reader can tell a leak from a genuine conflict without logging into the runner.
Acceptance Criteria
- A deliberately failed gate run leaves no GPU compute process behind: after the job completes,
nvidia-smi --query-compute-apps=pidandpgrep -f mlxcel-server-webuiboth return empty onlablup-dgxspark21, with the run ID recorded in the PR body. - The cleanup runs on the failure path and reports what it did, while the passing path still records
server_shutdown.exit_code == 0,forced: falseandprocess_group_empty: true, sosummarize_activity_evidence.py:build_summaryaccepts the summary unchanged. - A leaked orphan seeded by hand before a run is identified as a leak and named with its pid, command and deleted cwd, and is distinguished from the current run's own server rather than only counted.
-
terminate_ownedcannot raise out of its own SIGKILL fallback: a test inscripts/webui/verify_activity_performance_tests.pyusing a process that ignores SIGTERM and outlives the wait shows the function returning a result withforced: trueinstead of propagatingsubprocess.TimeoutExpired. A second test covers the escalation gap: a parent that exits inside the eight second wait while leaving a live process in its group must end with the group empty, which current code does not achieve because:393recordsprocess_group_emptywithout acting on it. Both tests fail on current code. - The gate refuses to score while the host is busy: with a synthetic load running on the runner, the step waits and then fails closed naming host load, and with the load removed the same commit passes.
- The full evidence JSON records the host state observed at gate start, and the summary still passes
scripts/webui/summarize_activity_evidence.py. - The gate still runs inside the real
WebUI installed artifactjob on every WebUI or Rust change, and the change is integrated into that job rather than left as a standalone script. - Everything above lands in a single PR, with the verification block run once at the end and pushes kept to a minimum, because every push touching this job costs a full GB10 gate cycle.
Deliberately not an acceptance criterion here: the three-way "no regression" / "unresolved within the measurement floor" / "regression" reporting, with paired_range_percent and baseline_cv_percent retained in the summary, belongs to #1925 and must not be duplicated in this PR.
Verification
python3 scripts/webui/verify_activity_performance_tests.py
python3 scripts/webui/summarize_activity_evidence_tests.py
make verify-webui-helper-tests
# On lablup-dgxspark21, immediately after a deliberately failed gate run:
nvidia-smi --query-compute-apps=pid,used_gpu_memory --format=csv
pgrep -af mlxcel-server-webui
A pass is: helper tests green; no GPU compute process and no mlxcel-server-webui left behind after the deliberately failed run; the quiet-host wait visible in the job log under synthetic load and the same commit green once the load is removed; and a green WebUI installed artifact run on an otherwise unchanged tree.
Technical Considerations
Related: #1925 (verdict rule, open and claimed; coordinate before touching the scoring path). lablup-dgxspark21 serves six ci.yml jobs plus the GB10 job in release.yml:521, so an orphaned server holding 4.8 GiB of GPU memory for hours blocks far more than this one gate, which is why the cleanup should fail loudly when it cannot reap rather than reap silently.
- Lingua principale
- Rust
- Stelle
- 471
- Fork
- 55
- Merge medio
- 9h 36m
- PR unite (30g)
- 289
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di lablup/mlxcel
-
area:core priority:low status:ready type:chore
Difficoltà 2/5 1-3 ore Idoneità per principianti 90/100
I maintainer di solito rispondono entro 1 giorno
-
priority:low status:ready type:docs
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 95/100
I maintainer di solito rispondono entro 1 giorno
-
docs(webpage): add a webpage/site README covering the pnpm/uv/zensical build and deploy contractApertapriority:low status:ready type:docs
Difficoltà 1/5 1-3 ore Idoneità per principianti 86/100
I maintainer di solito rispondono entro 1 giorno
-
priority:low status:ready type:docs
Difficoltà 2/5 1-3 ore Idoneità per principianti 92/100
I maintainer di solito rispondono entro 1 giorno
-
priority:medium status:ready type:docs
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di lablup/mlxcel
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
aws-samples/sample-pacer#76 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
axodotdev/cargo-dist#2523 ·
I maintainer di solito rispondono entro 2 giorni