🐛 After SIGTERM, `cloudflared tunnel run` uses 100% of one CPU core for the whole graceful-shutdown period
評価
調査の方向性
Issueは、supervisor/supervisor.go内でSupervisor.RunがgracefulShutdownCを対象にしているselectのbusy loopを指しています。そのループとwaitForSignalでのシグナル処理を読み、initialize()は影響を受けないことを確認してください。SIGTERM後、接続のdrain中にsupervisorがブロックし、トンネル接続が終了するか猶予期間が終了すると、引き続き終了することを検証してください。
索引モデルが issue の本文から書いたものです。
説明
Describe the bug
After SIGTERM, cloudflared tunnel run uses 100% of one CPU core for the whole graceful-shutdown period (up to --grace-period). It uses ~0% before the signal, including while it serves a long-lived request.
The cause is a busy loop in Supervisor.Run (supervisor/supervisor.go). On SIGTERM, waitForSignal calls close(graceShutdownC). The supervisor's main loop selects on that channel:
case <-s.gracefulShutdownC:
shuttingDown = true
A closed channel is always ready, so this case is selected on every iteration and the for { select { ... } } loop never blocks. The loop runs until the last tunnel connection exits or the process is killed.
To Reproduce
- Run a local origin that serves a long-lived response (e.g. an SSE endpoint that writes one event per second for 90s).
- Run
cloudflared tunnel --grace-period 100s --no-autoupdate --metrics 127.0.0.1:20241 run --token <token> - Open the long-lived request through the tunnel hostname.
kill -TERM <cloudflared pid>- Sample CPU (e.g.
/proc/<pid>/statutime+stime over 3s) or watchtop: cloudflared stays at 100% of one core until it exits. curl '127.0.0.1:20241/debug/pprof/goroutine?debug=2'shows the supervisor goroutine runnable inside theRunselect loop, while the rest are blocked.
If it's an issue with Cloudflare Tunnel:
4. Tunnel ID: f499dceb-781e-470b-9943-3c1aa7674c30
5. cloudflared config: no config file; remotely managed tunnel run with --token, --grace-period 100s, --no-autoupdate. Protocol: QUIC (default). The loop is in the supervisor, so it should not depend on the protocol.
Expected behavior
While draining, cloudflared should sit near-idle, blocked until in-flight requests finish or the grace period ends.
Environment and versions
- OS: Linux (Debian 13 trixie container, kernel 7.1.5)
- Architecture: AMD64 (x86_64)
- Version: 2026.9.3 (built 2026-09-24). The same code is on
master.
Logs and errors
CPU measured over 3s windows:
idle cpu%: 0
streaming cpu%: 0
drain1 cpu%: 100
drain2 cpu%: 100
drain3 cpu%: 101
Log at SIGTERM:
INF Initiating graceful shutdown due to signal terminated ...
INF Unregistered tunnel connection connIndex=2 event=0 ip=198.41.192.7
INF Unregistered tunnel connection connIndex=1 event=0 ip=198.41.192.57
Goroutine dump during drain (54 goroutines; only the pprof handler is running and this one is runnable):
goroutine 61 [runnable]:
context.(*cancelCtx).Done(...)
/usr/local/go/src/context/context.go:448
github.com/cloudflare/cloudflared/supervisor.(*Supervisor).Run(...)
.../cloudflared/supervisor/supervisor.go:137 +0x324
github.com/cloudflare/cloudflared/supervisor.StartTunnelDaemon(...)
.../cloudflared/supervisor/tunnel.go:104 +0x50
github.com/cloudflare/cloudflared/cmd/cloudflared/tunnel.StartServer.func5()
.../cloudflared/cmd/cloudflared/tunnel/cmd.go:521 +0x74
Additional context
Suggested fix: set the channel to nil after the first receive. A nil channel is never ready in a select, so the loop blocks again.
case <-s.gracefulShutdownC:
shuttingDown = true
s.gracefulShutdownC = nil
initialize() also selects on gracefulShutdownC, but it returns immediately on receive, so it is not affected.
Impact: we run cloudflared as the ingress next to our app server and drain for up to 100s on each deploy. During that time cloudflared takes a full core away from the app server while the app server finishes in-flight requests.
- 主要言語
- Go
- スター
- 16k
- フォーク
- 1.5k
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
cloudflare/cloudflared のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
cloudflare/cloudflared#1761 ·
-
🐛 cfRay is missing from proxied request logs (newHTTPLogger discards the zerolog context)対応中かも @pankajc46 が 1 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
cloudflare/cloudflared#1756 ·
-
tunnel route ip show: --filter-network-is-subset-of sends the superset filter対応中かも @wangyusheng1985 が 1 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 92/100
cloudflare/cloudflared#1750 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
cloudflare/cloudflared#1748 ·
-
🐛 QUIC Hijack() skips the status-written check that HTTP/2 enforces対応中かも @Asthenia0412 が 11 日前に担当しました。 オープンPriority: Normal Type: Bug
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
cloudflare/cloudflared#1747 ·
cloudflare/cloudflared の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
OpenTollGate/tollgate-module-basic-go#833 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
siyuan-note/siyuan#20353 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信
-
attributes-natural-language "en-US" is rejected by PAPPL >= 1.4.12 printers (RFC 8011 requires lowercase)対応中かも @ChrisEdgington が今日担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 84/100
OpenPrinting/ipp-usb#140 ·