Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

bug(sandbox): network broker logs expected socket races as warnings

オープン
#3,526 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
52/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
linux, rust

調査の方向性

Start by tracing openshell_sandbox::network_broker and its dispatch_notification error handling on the current main branch. Reproduce the notification-target and socket-race cases described in the issue, then inspect existing broker tests. Done means transient ENOTCONN/ESRCH cases no longer create warning storms, policy and fatal failures remain distinguishable, and tests cover disappearing targets and closing sockets.

索引モデルが issue の本文から書いたものです。

説明

area:sandbox os:linux topic:observability

This was generated by AI during triage.

Agent Diagnostic

  • Skills loaded: diagnose, OpenShell repository create-github-issue, and cluster inspection workflows
  • OpenShell version tested: 0.0.117-dev.211+g4cd5e5478
  • Latest release checked: v0.0.116; the issue reproduces on a newer development build containing the RFC 0012 sandbox runtime
  • Known fixes reviewed: release notes for v0.0.116, RFC 0012 implementation PR #2942, and current main network-broker handling
  • Possible duplicates reviewed: searched all OpenShell issues and merged PRs for the exact network notification denied text, Socket not connected, and related network/boundary warnings. No matching issue was found. #3396 concerns a boundary reconnect that becomes fatal; the sandboxes here remained healthy.
  • Findings: six healthy Kubernetes workload pods emitted 4,183 sandbox network notification denied warnings in 12 hours. Nearly all were syscall 52 returning ENOTCONN; the remainder were syscall 62 returning ESRCH. All pods stayed Ready with zero restarts and there were no error-level runtime logs. Current main logs every dispatch_notification error at WARN, regardless of whether it represents an enforcement decision or an expected process/socket race.
  • Remaining reason for filing: these per-syscall warnings are high-volume and misleading, and obscure actionable policy or boundary failures.

Description

Actual behavior: The RFC 0012 sandbox network broker emits a WARN for every failed seccomp notification dispatch:

WARN openshell_sandbox::network_broker: sandbox network notification denied (tid=..., syscall=52): Socket not connected (os error 107)
WARN openshell_sandbox::network_broker: sandbox network notification denied (tid=..., syscall=62): No such process (os error 3)

During a 12-hour observation of six active agent workloads, this produced 4,183 network-broker warnings. The workloads remained Ready, had zero restarts, and recorded no error-level runtime events. A two-hour sample was dominated by 910 ENOTCONN warnings and 34 ESRCH warnings.

The current handler groups all dispatch_notification errors under the message sandbox network notification denied, even when the returned errno describes a socket/process race rather than an OpenShell policy denial. This makes benign application behavior look like a security or isolation failure and makes genuine broker failures difficult to find.

Expected behavior: Expected transient process/socket races should not create one warning per syscall. They should be handled at debug/trace level, aggregated, or rate-limited. Genuine OpenShell policy denials and broker-health failures must remain clearly observable and distinguishable from application-originated errnos.

Reproduction Steps

  1. Deploy the RFC 0012 sandbox runtime on Linux with seccomp network mediation enabled.
  2. Run a long-lived agent workload that performs ordinary concurrent HTTP/socket activity and periodically starts and exits subprocesses.
  3. Collect the workload pod's runtime logs for several polling cycles.
  4. Count sandbox network notification denied messages.
  5. Observe repeated syscall 52/ENOTCONN and syscall 62/ESRCH warnings while the sandbox remains Ready and functional.

Environment

  • OS: Ubuntu 24.04, Linux amd64
  • Kubernetes: v1.35.7-gke.1222000
  • Agent Sandbox controller: v0.5.0, v1beta1 API
  • Docker: not applicable; Kubernetes compute driver
  • OpenShell: 0.0.117-dev.211+g4cd5e5478
  • Deployment: Helm 0.0.0-dev, RFC 0012 separate workload and supervisor pods
  • Latest release checked: yes, v0.0.116; the tested development build is newer
  • Possible duplicates checked: yes; #3396 is related to fatal boundary recovery but does not cover this non-fatal per-syscall warning storm

Logs

# Counts by workload over 12 hours:
workload-1 network_denied=190 errors=0 restarts=0
workload-2 network_denied=90  errors=0 restarts=0
workload-3 network_denied=1224 errors=0 restarts=0
workload-4 network_denied=897 errors=0 restarts=0
workload-5 network_denied=1046 errors=0 restarts=0
workload-6 network_denied=736 errors=0 restarts=0

# Dominant messages over a two-hour sample:
910 WARN ... syscall=52: Socket not connected (os error 107)
 34 WARN ... syscall=62: No such process (os error 3)

Acceptance Criteria

  • Expected ENOTCONN and ESRCH process/socket races do not emit an unbounded WARN per syscall.
  • Genuine policy denials remain observable as explicit policy decisions rather than being conflated with application errno handling.
  • Fatal listener or broker-health failures remain at error level and retain actionable context.
  • Telemetry provides aggregate counts or suitably rate-limited diagnostics for suppressed transient failures.
  • Tests cover notification targets disappearing and sockets closing between notification receipt and dispatch without producing warning storms.
主要言語
Rust
スター
8.7k
フォーク
1.3k
平均マージ
2日 6時間
マージ済み PR(30日)
301

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/OpenShell のほかの issue

NVIDIA/OpenShell の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。