bug: surface supervisor startup failures to sandbox commands

オープン
#3,519 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
52/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
rust
領域
backend, cli

調査の方向性

Start by tracing the sandbox create/run command result through supervisor startup and readiness handling, using the reproduced /bin/cat permission failure as the test case. Identify where ContainerExited is produced and where startup diagnostics can be classified and sanitized. Done means the CLI reports a distinct startup failure and a unit or integration test covers entrypoint execution failure without exposing raw logs or credentials.

索引モデルが issue の本文から書いたものです。

説明

User Story

As an OpenShell operator, I want a sandbox command to report a meaningful startup failure so that I can diagnose failed sandbox launches without manually inspecting runtime logs.

Problem Statement

When a supervisor cannot start the workload entrypoint, the sandbox lifecycle is reduced to a generic container exit. For example, a rootless Podman run that fails to execute /bin/cat with Permission denied (os error 13) is reported by the CLI as ContainerExited with exit code 1. The useful error is only present in the supervisor container stderr; the workload log instead records the subsequent control-channel termination.

Impact / Why This Matters

Operators cannot distinguish an entrypoint execution failure from an ordinary workload exit through the OpenShell command result. The current workaround is to preserve failed containers and inspect Podman logs manually, which is runtime-specific, slow, and unsuitable for automated diagnostics.

Acceptance Criteria

  • A supervisor failure before sandbox readiness is represented as a distinct sandbox startup failure rather than only ContainerExited.
  • openshell sandbox create or run exposes a concise, sanitized diagnostic that identifies the failed startup stage and failure class.
  • The implementation does not propagate arbitrary workload or supervisor log output, credentials, or environment values into lifecycle status.
  • The behavior is covered by a unit or integration test for an entrypoint execution failure.

Reproduction Steps

  1. Configure a rootless Podman gateway with the current restrictive sandbox policy.
  2. Create a sandbox from Alpine with an explicit executable entrypoint, for example -- /bin/cat /proc/self/uid_map.
  3. Observe the command fail with a generic container-exited status.
  4. Inspect the supervisor container logs to find the underlying Permission denied (os error 13) spawn failure.

Environment

  • OpenShell: current main / Alpine-default work
  • OS: Fedora tmachine VM
  • Runtime: rootless Podman

Logs

Error: process error: boundary process leaf: start process supervisor leaf:
  spawn delegated workload process
    failed to spawn sandbox entrypoint process '/bin/cat'
    Permission denied (os error 13)

Related Work

The Alpine-default work is being tracked in PR #3386. This issue is deliberately scoped to reporting the failure; the policy/entrypoint failure itself can be addressed independently.

主要言語
Rust
スター
8.7k
フォーク
1.3k
平均マージ
2日 6時間
マージ済み PR(30日)
236

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/OpenShell のほかの issue

NVIDIA/OpenShell の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。