Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Zombie connection after lightningd ignores `connectd_peer_spoke`

オープン
#9,369 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

@morehouse がすでに取り組んでいます。

2026年8月2日 から。

  • #9370 @morehouse による — オープン

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
55/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
c
領域
networking

調査の方向性

connectd/multiplex.c と lightningd/peer_control.c の関連する処理から始め、特に handle_peer_spoke が応答しない 6 つのパスを確認します。提案されている connectd_peer_no_subd メッセージを両方のコンポーネントで追跡します。完了条件は、subdaemon の起動または検索に失敗した場合に対応する subd が解放され、connectd が peer メッセージの読み取りを再開することです。

索引モデルが issue の本文から書いたものです。

説明

When connectd reads a peer message whose channel_id has no attached subd, it creates a subd with conn == NULL and sends connectd_peer_spoke to lightningd, asking it to start up a subdaemon and respond with connectd_peer_connect_subd and the file descriptor that should be assigned to conn.

While connectd is waiting for lightningd's response, it stops reading all messages from the peer:

https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/connectd/multiplex.c#L1506-L1538

The io_wait at the end is only released when either lightningd responds, or when CLN needs to send a message to that peer (e.g., gossip flush). If neither happens, the connection becomes a zombie and all peer messages are ignored until eventually the ping timeout expires and the connection is dropped.

There are currently six situations where lightningd's handle_peer_spoke never responds:

  1. The peer sends an error for its first message (e.g., data loss recovery).
    https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/lightningd/peer_control.c#L2067-L2072

  2. channeld dies from a bug or protocol violation, and the peer sends another message for that channel before lightningd recognizes the channeld died. An alternative (benign) way to trigger this is when lightningd has already created the requested subdaemon but connectd hasn't processed it yet.
    https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/lightningd/peer_control.c#L2074-L2081

  3. The peer attempts a channel_reestablish while the node is shutting down.
    https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/lightningd/peer_control.c#L2083-L2097

  4. The peer sends a message after channeld was killed without a status message (e.g., OOM).
    https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/lightningd/peer_control.c#L2100-L2109

  5. openingd fails to spawn (e.g., fd limit reached).
    https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/lightningd/peer_control.c#L2143-L2145

  6. dualopend fails to spawn (e.g., fd limit reached).
    https://github.com/ElementsProject/lightning/blob/ae53e8775e57ddf8debfd6fc7b8e3cb5e2c93dc6/lightningd/peer_control.c#L2163-L2165

Suggested fix

Add a new message connectd_peer_no_subd that lightningd can respond with when handle_peer_spoke fails in the above situations. Then connectd knows to free the matching subd and continue reading from the connection.

Discovery

This bug was discovered while fuzzing CLN with smite. Smite would send a channel_ready message with an incorrect channel_id, causing channeld to exit. Then smite would send another channel message, which would trigger the Situation 2 race condition and zombify the connection.

主要言語
C
スター
3.1k
フォーク
1k
平均マージ
3日 10時間
マージ済み PR(30日)
40

環境構築

  • Dockerfile または Docker Compose ファイルあり
  • プルリクエストのテンプレートあり
  • コントリビューションガイドなし

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

ElementsProject/lightning のほかの issue

ElementsProject/lightning の issue をすべて見る

似ている issue

C の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。