Orchestrator GRPC_PORT takes 20-30 min to bind on startup when /run/netns has many residual ns-* from a previously drained node
メンテナーはふだん 3 日以内に返信
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 38/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- go, grpc
調査の方向性
packages/orchestrator/pkg/factories/run.go から始め、startupreclaim.Run、NewStorageLocal、pool の生成、745-974 行付近の cmux の bind を追跡します。packages/orchestrator/pkg/sandbox/network/reclaim.go と packages/orchestrator/pkg/service/info.go を読み、その後、関連する orchestrator の起動テストを実行するか、残存する ns-* エントリで再現します。合意した起動設計によって health の動作を正しく保ちつつ、GRPC_PORT がチェックを受け付けるまでの 20~30 分間を回避できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Summary
On a host that previously ran a large number of sandboxes and was then drained to 0 sandboxes, a lot of ns-* network namespaces can remain in /run/netns (leaked slots from failed/partial teardown). When we upgrade/restart the orchestrator on such a host, the gRPC service (GRPC_PORT) does not bind for 20–30 minutes. During that window the port is closed → Nomad health checks get connection-refused → the node looks dead / restart-flaps.
Environment
- Deployment: self-hosted, Nomad
type = "system"job,GRPC_PORTbound as a static port - Orchestrator running the sandbox runtime (
ORCHESTRATOR_SERVICESincludesorchestrator) - Host state before upgrade: 0 running sandboxes, but a large number of residual
ns-*entries in/run/netns
Root cause (from reading the code)
The startup path is strictly serial, and the whole reclaim chain runs before the listener is bound:
acquireOrchestratorLock(flock) —packages/orchestrator/pkg/factories/run.go:398startupreclaim.Run(...)—run.go:745- →
network.ReclaimLeakedSlots(netnsDir, ...)—packages/orchestrator/pkg/sandbox/network/reclaim.go:15 network.NewStorageLocal(...)(must run after reclaim so leftoverns-*aren't snapshotted asforeign) —run.go:767, see note atrun.go:736-738- network pool
Populate(...) - … only later:
cmuxbindsGRPC_PORT—run.go:974
The bottleneck is step 3. ReclaimLeakedSlots iterates leaked slots one by one and calls slot.RemoveNetwork() for each:
// packages/orchestrator/pkg/sandbox/network/reclaim.go
for _, idx := range slots {
slot, err := NewSlot(fmt.Sprintf("startup-reclaim-%d", idx), idx, config, egressProxy)
...
if err := slot.RemoveNetwork(); err != nil { ... } // netns + netlink + iptables/nftables teardown per slot
}
With thousands of residual namespaces, this serial per-namespace teardown takes tens of minutes, and because it sits ahead of cmux.Serve(), GRPC_PORT stays unbound the entire time.
Impact
- Node is unreachable on
GRPC_PORTfor 20–30 min after every upgrade/restart on a "dirty" host. - With Nomad
system+ static port, health checks fail with connection-refused (not 503), so it reads as failed rather than starting → restart churn. - The more sandboxes the host handled historically, the worse the leak, the longer the outage.
Proposed solutions (in rough priority order)
Solution A — Bind the port + serve /health early, run reclaim/populate in the background
Decouple "port is listening + health endpoint answering" from "sandbox runtime fully initialized":
- Keep
acquireOrchestratorLockfirst and blocking (single-instance guard; reclaim mutates host-level netns/iptables and must be exclusive). - Bind
cmuxand start the HTTP/healthserver immediately, reporting not healthy (503) while initializing. - Run reclaim →
NewStorageLocal→ pool populate →server.New(...)→ gRPCRegisterServicein a background goroutine, then callgrpcServer.Serve(grpcListener)and flip status toHealthy.
Two required correctness points:
packages/orchestrator/pkg/service/info.go:94currently initializes status toHealthy. It must start as a non-healthy state (e.g.Starting/Standby), otherwise the early/healthreturns 200 and Nomad routes gRPC traffic before services are registered.- gRPC forbids
RegisterServiceafterServe(); theServe(grpcListener)call must happen only after allRegister*complete (early gRPC connections buffer in the cmux matcher until then).
Effect: Nomad sees an open port returning 503 → treats the node as starting, not failed → no restart flap; the long reclaim no longer blocks bind.
Solution B — Parallelize ReclaimLeakedSlots
The per-slot teardowns are independent. Bounding the loop with a worker pool (e.g. errgroup with a concurrency limit) would cut the reclaim wall-time roughly linearly. This helps even without Solution A. Note: the intra-chain ordering reclaim → NewStorageLocal must be preserved; only the loop inside reclaim is parallelized.
Solution C — Make startup reclaim bounded / deferrable
- A time or count budget for startup reclaim, deferring the remainder to a background reconciler after the node is already serving.
- Or an opt-in async-reclaim mode (
DisableStartupReclaimalready exists as a related knob).
Additional questions for maintainers
- Is such a large
ns-*leak on drain expected, or does it point to a teardown path that should have cleaned these during normal operation? Fixing the leak source would reduce reclaim load in the first place. - Would you accept Solution A (early-bind + background init) as the primary fix, with Solution B as a complementary speedup?
I'm happy to open a PR for Solution A and/or Solution B if the direction is agreeable.
- 主要言語
- Go
- スター
- 1.7k
- フォーク
- 468
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
e2b-dev/runtime のほかの issue
-
[Bug]: flock() on a mounted volume hangs forever (mount is missing `nolock`)対応中かも @AdaAibaby が 34 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 3 日以内に返信
-
sandbox cache: StartRemoving state transition not broadcast, all allocations see stale Running state対応中かも @AdaAibaby が 50 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
メンテナーはふだん 3 日以内に返信
-
kernel: ipv6.disable=0 with no IPv6 routing causes Happy Eyeballs latency on all outbound sandbox connections対応中かも @AdaAibaby が 54 日前に担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 86/100
メンテナーはふだん 3 日以内に返信
-
fix(envd): malformed error message %!w(<nil>) when watch path is not a directory対応中かも @chill-czar が 54 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
メンテナーはふだん 3 日以内に返信
-
fix(api): POST /sandboxes/{id}/connect accepts non-positive timeout values causing immediate termination対応中かも @chill-czar が 54 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 3 日以内に返信
e2b-dev/runtime の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 75/100
slavakurilyak/awesome-ai-agents#742 ·
メンテナーはふだん 1 日以内に返信
-
bug go
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
genkit-ai/genkit#6761 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 87/100
メンテナーはふだん 2 日以内に返信