CI: ASan/Valgrind shards cancelled mid-run — runner receives shutdown signal (likely OOM on 16GB hosted runner)
メンテナーはふだん 2 日以内に返信
関連するプルリクエストがすでにマージされています。
- #9280 @nGoline による — マージ済み
評価
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 68/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 静か
- 技術スタック
- c, github-actions
- 領域
- ci-cd, performance, testing-qa
調査の方向性
.github/workflows/ci.yaml の 665、729、771、773 行付近にある sanitizer ジョブと Valgrind ジョブから始め、matrix 設定、worker 数、テストコマンドを確認します。影響を受ける shard を再現し、選択したリソース調整によって runner のシャットダウンによるキャンセルが防止され、同じ階層の shard が引き続き正常に実行されることを確認します。
索引モデルが issue の本文から書いたものです。
説明
Summary
The heaviest CI jobs (ASan/UBSan and Valgrind Test CLN) intermittently get their runner killed part-way through the test step. The GitHub annotation only shows Error: The operation was canceled, with no test failure, which makes it look like "almost every test broke". On the ASan/UBSan (3/6) shard it reproduces at almost exactly the 10% mark.
This is not a test failure and not the concurrency/fail-fast config. The runner VM itself is being terminated mid-run, most likely from memory exhaustion on the 16GB hosted runner.
Evidence
The real cause is one line above the cancellation in the raw job log (hidden from the annotations view):
##[error]The runner has received a shutdown signal. This can happen when the runner service is stopped, or a manually started runner is canceled.
[gw2] [ 10%] PASSED tests/test_closing.py::test_segwit_shutdown_script
##[error]The operation was canceled.
No test errored. The runner agent died, so every in-flight and pending test on that shard is reported cancelled.
Sample runs:
- master push: https://github.com/ElementsProject/lightning/actions/runs/28660219314/job/85000852219
- PR #9143: https://github.com/ElementsProject/lightning/actions/runs/28660928500/job/85003365800
What it is not
-
Not the
concurrency/cancel-in-progressgroup (.github/workflows/ci.yaml:9). That cancels the whole run. In the master run above only three shards died while every sibling passed:ASan/UBSan (1/6) success ASan/UBSan (3/6) FAILURE ASan/UBSan (4/6) success ... Valgrind (10/12) FAILURE Valgrind (12/12) FAILURE (other 10 succeeded) -
Not
fail-fast. Bothintegration-sanitizersandintegration-valgrindsetfail-fast: false(ci.yaml:729,ci.yaml:665), so a sibling cannot cancel them. -
Not a timeout. The shard died at ~6 min; the job timeout is 120 min and the per-test timeout is 1800s.
Individual runner VMs are dying independently, concentrated on exactly the two most memory-heavy job types.
Likely root cause: memory exhaustion on the 16GB hosted runner
runs-on: ubuntu-24.04is a standard GitHub-hosted runner: 4 vCPU / 16GB RAM.- The test step runs
pytest -n $(($(nproc) + 1))-> 5 workers (ci.yaml:773; log confirmscreated: 5/5 workers). - The build under test is
compile-clang-sanitizers; ASan roughly triples RSS and Valgrind is heavier still. - At the moment of death,
gw0had been runningtests/test_askrene.py::test_real_data(loads the real mainnet gossip map, one of the most memory-hungry tests) for ~3 min alongside four other worker node-clusters.
Five parallel ASan lightningd+bitcoind clusters with test_real_data in the mix peaks past 16GB, and the host force-terminates the VM.
Why always ~10% on ASan/UBSan (3/6): --test-group-random-seed=42 (ci.yaml:771) fixes the shard's test set and order deterministically, so the shard hits its memory peak at the same point every run. Other runs hit the same wall on different heavy shards (e.g. Valgrind 10/12, 12/12).
The one thing not provable from the outside is OOM vs. a random hosted-runner reclaim, because GitHub does not expose the VM dmesg. The concentration on only ASan/Valgrind plus the deterministic repro point argues strongly for resource exhaustion rather than random reclaim.
Suggested fixes (cheapest first)
- Drop the worker count on the sanitizer/valgrind test steps: change
-n $(($(nproc) + 1))to-n $(nproc)or-n $(($(nproc) - 1))for those two jobs only. Fewer concurrent node-clusters means lower peak memory. - Cap ASan memory via
ASAN_OPTIONS=quarantine_size_mb=64:malloc_context_size=5in the sanitizer job env. - Add an
if: failure()step runningdmesg | grep -i oom || journalctl -k | tailso the next failure records the actual cause.
Likely the same underlying overload as the timing-based flakes such as #9268 (an 11s delay to activate connectd under load). Related tracking issue: #9222.
- 主要言語
- C
- スター
- 3.1k
- フォーク
- 1k
- 平均マージ
- 3日 10時間
- マージ済み PR(30日)
- 40
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
ElementsProject/lightning のほかの issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
ElementsProject/lightning#9593 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
ElementsProject/lightning#9322 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
ElementsProject/lightning#9206 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
ElementsProject/lightning#9187 · コメント 1 件 · リアクション 1 件 ·
メンテナーはふだん 2 日以内に返信
-
QA
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
ElementsProject/lightning#9117 · コメント 2 件 ·
メンテナーはふだん 2 日以内に返信
ElementsProject/lightning の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
kovidgoyal/kitty#10625 ·
メンテナーはふだん 1 日以内に返信
-
Feature Status: Needs Triage
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
メンテナーはふだん 1 日以内に返信
-
docs
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
RubyMetric/chsrc#396 ·