Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

bug(sandbox): network broker logs expected socket races as warnings

未关闭
#3,526 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
52/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
linux, rust

调研方向

Start by tracing openshell_sandbox::network_broker and its dispatch_notification error handling on the current main branch. Reproduce the notification-target and socket-race cases described in the issue, then inspect existing broker tests. Done means transient ENOTCONN/ESRCH cases no longer create warning storms, policy and fatal failures remain distinguishable, and tests cover disappearing targets and closing sockets.

由索引模型根据 Issue 内容生成。

描述

area:sandbox os:linux topic:observability

This was generated by AI during triage.

Agent Diagnostic

  • Skills loaded: diagnose, OpenShell repository create-github-issue, and cluster inspection workflows
  • OpenShell version tested: 0.0.117-dev.211+g4cd5e5478
  • Latest release checked: v0.0.116; the issue reproduces on a newer development build containing the RFC 0012 sandbox runtime
  • Known fixes reviewed: release notes for v0.0.116, RFC 0012 implementation PR #2942, and current main network-broker handling
  • Possible duplicates reviewed: searched all OpenShell issues and merged PRs for the exact network notification denied text, Socket not connected, and related network/boundary warnings. No matching issue was found. #3396 concerns a boundary reconnect that becomes fatal; the sandboxes here remained healthy.
  • Findings: six healthy Kubernetes workload pods emitted 4,183 sandbox network notification denied warnings in 12 hours. Nearly all were syscall 52 returning ENOTCONN; the remainder were syscall 62 returning ESRCH. All pods stayed Ready with zero restarts and there were no error-level runtime logs. Current main logs every dispatch_notification error at WARN, regardless of whether it represents an enforcement decision or an expected process/socket race.
  • Remaining reason for filing: these per-syscall warnings are high-volume and misleading, and obscure actionable policy or boundary failures.

Description

Actual behavior: The RFC 0012 sandbox network broker emits a WARN for every failed seccomp notification dispatch:

WARN openshell_sandbox::network_broker: sandbox network notification denied (tid=..., syscall=52): Socket not connected (os error 107)
WARN openshell_sandbox::network_broker: sandbox network notification denied (tid=..., syscall=62): No such process (os error 3)

During a 12-hour observation of six active agent workloads, this produced 4,183 network-broker warnings. The workloads remained Ready, had zero restarts, and recorded no error-level runtime events. A two-hour sample was dominated by 910 ENOTCONN warnings and 34 ESRCH warnings.

The current handler groups all dispatch_notification errors under the message sandbox network notification denied, even when the returned errno describes a socket/process race rather than an OpenShell policy denial. This makes benign application behavior look like a security or isolation failure and makes genuine broker failures difficult to find.

Expected behavior: Expected transient process/socket races should not create one warning per syscall. They should be handled at debug/trace level, aggregated, or rate-limited. Genuine OpenShell policy denials and broker-health failures must remain clearly observable and distinguishable from application-originated errnos.

Reproduction Steps

  1. Deploy the RFC 0012 sandbox runtime on Linux with seccomp network mediation enabled.
  2. Run a long-lived agent workload that performs ordinary concurrent HTTP/socket activity and periodically starts and exits subprocesses.
  3. Collect the workload pod's runtime logs for several polling cycles.
  4. Count sandbox network notification denied messages.
  5. Observe repeated syscall 52/ENOTCONN and syscall 62/ESRCH warnings while the sandbox remains Ready and functional.

Environment

  • OS: Ubuntu 24.04, Linux amd64
  • Kubernetes: v1.35.7-gke.1222000
  • Agent Sandbox controller: v0.5.0, v1beta1 API
  • Docker: not applicable; Kubernetes compute driver
  • OpenShell: 0.0.117-dev.211+g4cd5e5478
  • Deployment: Helm 0.0.0-dev, RFC 0012 separate workload and supervisor pods
  • Latest release checked: yes, v0.0.116; the tested development build is newer
  • Possible duplicates checked: yes; #3396 is related to fatal boundary recovery but does not cover this non-fatal per-syscall warning storm

Logs

# Counts by workload over 12 hours:
workload-1 network_denied=190 errors=0 restarts=0
workload-2 network_denied=90  errors=0 restarts=0
workload-3 network_denied=1224 errors=0 restarts=0
workload-4 network_denied=897 errors=0 restarts=0
workload-5 network_denied=1046 errors=0 restarts=0
workload-6 network_denied=736 errors=0 restarts=0

# Dominant messages over a two-hour sample:
910 WARN ... syscall=52: Socket not connected (os error 107)
 34 WARN ... syscall=62: No such process (os error 3)

Acceptance Criteria

  • Expected ENOTCONN and ESRCH process/socket races do not emit an unbounded WARN per syscall.
  • Genuine policy denials remain observable as explicit policy decisions rather than being conflated with application errno handling.
  • Fatal listener or broker-health failures remain at error level and retain actionable context.
  • Telemetry provides aggregate counts or suitably rate-limited diagnostics for suppressed transient failures.
  • Tests cover notification targets disappearing and sockets closing between notification receipt and dispatch without producing warning storms.
主要语言
Rust
星标
8.7k
派生
1.3k
平均合并
2 天 6 小时
30 天内合并 PR
301

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/OpenShell 的其他 Issue

查看 NVIDIA/OpenShell 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。