Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Orchestrator GRPC_PORT takes 20-30 min to bind on startup when /run/netns has many residual ns-* from a previously drained node

オープン
#3,615 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 3 日以内に返信

@AdaAibaby がすでに取り組んでいます。

2026年9月3日 から。

  • #3616 @AdaAibaby による — オープン

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
38/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
go, grpc

調査の方向性

packages/orchestrator/pkg/factories/run.go から始め、startupreclaim.Run、NewStorageLocal、pool の生成、745-974 行付近の cmux の bind を追跡します。packages/orchestrator/pkg/sandbox/network/reclaim.go と packages/orchestrator/pkg/service/info.go を読み、その後、関連する orchestrator の起動テストを実行するか、残存する ns-* エントリで再現します。合意した起動設計によって health の動作を正しく保ちつつ、GRPC_PORT がチェックを受け付けるまでの 20~30 分間を回避できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Summary

On a host that previously ran a large number of sandboxes and was then drained to 0 sandboxes, a lot of ns-* network namespaces can remain in /run/netns (leaked slots from failed/partial teardown). When we upgrade/restart the orchestrator on such a host, the gRPC service (GRPC_PORT) does not bind for 20–30 minutes. During that window the port is closed → Nomad health checks get connection-refused → the node looks dead / restart-flaps.

Environment

  • Deployment: self-hosted, Nomad type = "system" job, GRPC_PORT bound as a static port
  • Orchestrator running the sandbox runtime (ORCHESTRATOR_SERVICES includes orchestrator)
  • Host state before upgrade: 0 running sandboxes, but a large number of residual ns-* entries in /run/netns

Root cause (from reading the code)

The startup path is strictly serial, and the whole reclaim chain runs before the listener is bound:

  1. acquireOrchestratorLock (flock) — packages/orchestrator/pkg/factories/run.go:398
  2. startupreclaim.Run(...) — run.go:745
  3. → network.ReclaimLeakedSlots(netnsDir, ...) — packages/orchestrator/pkg/sandbox/network/reclaim.go:15
  4. network.NewStorageLocal(...) (must run after reclaim so leftover ns-* aren't snapshotted as foreign) — run.go:767, see note at run.go:736-738
  5. network pool Populate(...)
  6. … only later: cmux binds GRPC_PORT — run.go:974

The bottleneck is step 3. ReclaimLeakedSlots iterates leaked slots one by one and calls slot.RemoveNetwork() for each:

// packages/orchestrator/pkg/sandbox/network/reclaim.go
for _, idx := range slots {
    slot, err := NewSlot(fmt.Sprintf("startup-reclaim-%d", idx), idx, config, egressProxy)
    ...
    if err := slot.RemoveNetwork(); err != nil { ... }  // netns + netlink + iptables/nftables teardown per slot
}

With thousands of residual namespaces, this serial per-namespace teardown takes tens of minutes, and because it sits ahead of cmux.Serve(), GRPC_PORT stays unbound the entire time.

Impact

  • Node is unreachable on GRPC_PORT for 20–30 min after every upgrade/restart on a "dirty" host.
  • With Nomad system + static port, health checks fail with connection-refused (not 503), so it reads as failed rather than starting → restart churn.
  • The more sandboxes the host handled historically, the worse the leak, the longer the outage.

Proposed solutions (in rough priority order)

Solution A — Bind the port + serve /health early, run reclaim/populate in the background

Decouple "port is listening + health endpoint answering" from "sandbox runtime fully initialized":

  • Keep acquireOrchestratorLock first and blocking (single-instance guard; reclaim mutates host-level netns/iptables and must be exclusive).
  • Bind cmux and start the HTTP /health server immediately, reporting not healthy (503) while initializing.
  • Run reclaim → NewStorageLocal → pool populate → server.New(...) → gRPC RegisterService in a background goroutine, then call grpcServer.Serve(grpcListener) and flip status to Healthy.

Two required correctness points:

  • packages/orchestrator/pkg/service/info.go:94 currently initializes status to Healthy. It must start as a non-healthy state (e.g. Starting/Standby), otherwise the early /health returns 200 and Nomad routes gRPC traffic before services are registered.
  • gRPC forbids RegisterService after Serve(); the Serve(grpcListener) call must happen only after all Register* complete (early gRPC connections buffer in the cmux matcher until then).

Effect: Nomad sees an open port returning 503 → treats the node as starting, not failed → no restart flap; the long reclaim no longer blocks bind.

Solution B — Parallelize ReclaimLeakedSlots

The per-slot teardowns are independent. Bounding the loop with a worker pool (e.g. errgroup with a concurrency limit) would cut the reclaim wall-time roughly linearly. This helps even without Solution A. Note: the intra-chain ordering reclaim → NewStorageLocal must be preserved; only the loop inside reclaim is parallelized.

Solution C — Make startup reclaim bounded / deferrable
  • A time or count budget for startup reclaim, deferring the remainder to a background reconciler after the node is already serving.
  • Or an opt-in async-reclaim mode (DisableStartupReclaim already exists as a related knob).

Additional questions for maintainers

  • Is such a large ns-* leak on drain expected, or does it point to a teardown path that should have cleaned these during normal operation? Fixing the leak source would reduce reclaim load in the first place.
  • Would you accept Solution A (early-bind + background init) as the primary fix, with Solution B as a complementary speedup?

I'm happy to open a PR for Solution A and/or Solution B if the direction is agreeable.

主要言語
Go
スター
1.7k
フォーク
468
PR マージ指標
30日以内にマージされた PR はありません

環境構築

Codespaces で開く

このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

e2b-dev/runtime のほかの issue

e2b-dev/runtime の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。