Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

CI: ASan/Valgrind shards cancelled mid-run — runner receives shutdown signal (likely OOM on 16GB hosted runner)

オープン
#9,279 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

関連するプルリクエストがすでにマージされています。

  • #9280 @nGoline による — マージ済み

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
68/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
c, github-actions

調査の方向性

.github/workflows/ci.yaml の 665、729、771、773 行付近にある sanitizer ジョブと Valgrind ジョブから始め、matrix 設定、worker 数、テストコマンドを確認します。影響を受ける shard を再現し、選択したリソース調整によって runner のシャットダウンによるキャンセルが防止され、同じ階層の shard が引き続き正常に実行されることを確認します。

索引モデルが issue の本文から書いたものです。

説明

Summary

The heaviest CI jobs (ASan/UBSan and Valgrind Test CLN) intermittently get their runner killed part-way through the test step. The GitHub annotation only shows Error: The operation was canceled, with no test failure, which makes it look like "almost every test broke". On the ASan/UBSan (3/6) shard it reproduces at almost exactly the 10% mark.

This is not a test failure and not the concurrency/fail-fast config. The runner VM itself is being terminated mid-run, most likely from memory exhaustion on the 16GB hosted runner.

Evidence

The real cause is one line above the cancellation in the raw job log (hidden from the annotations view):

##[error]The runner has received a shutdown signal. This can happen when the runner service is stopped, or a manually started runner is canceled.
[gw2] [ 10%] PASSED tests/test_closing.py::test_segwit_shutdown_script
##[error]The operation was canceled.

No test errored. The runner agent died, so every in-flight and pending test on that shard is reported cancelled.

Sample runs:

What it is not

  • Not the concurrency / cancel-in-progress group (.github/workflows/ci.yaml:9). That cancels the whole run. In the master run above only three shards died while every sibling passed:

    ASan/UBSan (1/6) success   ASan/UBSan (3/6) FAILURE   ASan/UBSan (4/6) success ...
    Valgrind (10/12) FAILURE   Valgrind (12/12) FAILURE   (other 10 succeeded)
    
  • Not fail-fast. Both integration-sanitizers and integration-valgrind set fail-fast: false (ci.yaml:729, ci.yaml:665), so a sibling cannot cancel them.

  • Not a timeout. The shard died at ~6 min; the job timeout is 120 min and the per-test timeout is 1800s.

Individual runner VMs are dying independently, concentrated on exactly the two most memory-heavy job types.

Likely root cause: memory exhaustion on the 16GB hosted runner

  • runs-on: ubuntu-24.04 is a standard GitHub-hosted runner: 4 vCPU / 16GB RAM.
  • The test step runs pytest -n $(($(nproc) + 1)) -> 5 workers (ci.yaml:773; log confirms created: 5/5 workers).
  • The build under test is compile-clang-sanitizers; ASan roughly triples RSS and Valgrind is heavier still.
  • At the moment of death, gw0 had been running tests/test_askrene.py::test_real_data (loads the real mainnet gossip map, one of the most memory-hungry tests) for ~3 min alongside four other worker node-clusters.

Five parallel ASan lightningd+bitcoind clusters with test_real_data in the mix peaks past 16GB, and the host force-terminates the VM.

Why always ~10% on ASan/UBSan (3/6): --test-group-random-seed=42 (ci.yaml:771) fixes the shard's test set and order deterministically, so the shard hits its memory peak at the same point every run. Other runs hit the same wall on different heavy shards (e.g. Valgrind 10/12, 12/12).

The one thing not provable from the outside is OOM vs. a random hosted-runner reclaim, because GitHub does not expose the VM dmesg. The concentration on only ASan/Valgrind plus the deterministic repro point argues strongly for resource exhaustion rather than random reclaim.

Suggested fixes (cheapest first)

  1. Drop the worker count on the sanitizer/valgrind test steps: change -n $(($(nproc) + 1)) to -n $(nproc) or -n $(($(nproc) - 1)) for those two jobs only. Fewer concurrent node-clusters means lower peak memory.
  2. Cap ASan memory via ASAN_OPTIONS=quarantine_size_mb=64:malloc_context_size=5 in the sanitizer job env.
  3. Add an if: failure() step running dmesg | grep -i oom || journalctl -k | tail so the next failure records the actual cause.

Likely the same underlying overload as the timing-based flakes such as #9268 (an 11s delay to activate connectd under load). Related tracking issue: #9222.

主要言語
C
スター
3.1k
フォーク
1k
平均マージ
3日 10時間
マージ済み PR(30日)
40

環境構築

  • Dockerfile または Docker Compose ファイルあり
  • プルリクエストのテンプレートあり
  • コントリビューションガイドなし

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

ElementsProject/lightning のほかの issue

ElementsProject/lightning の issue をすべて見る

似ている issue

C の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。