Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

bug: distinguish intentional signal stops from runtime restarts

Aberta
#3,083 2 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
5/5
Tempo estimado
Mais de uma semana
Facilidade para iniciantes
45/100
Tipo de issue
Bug
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
docker, rust

Direção de pesquisa

Rastreie o tratamento dos códigos de saída 137/143 pelos drivers Docker e Podman e, em seguida, acompanhe como a intenção de parada do gateway e os snapshots atrasados do watcher atualizam o status do sandbox. Revise o comportamento de recuperação existente relacionado aos issues #2855 e #2179 e execute ou amplie a cobertura de regressão para ambos os drivers. Considera-se concluído quando as paradas intencionais continuarem sendo terminais e distinguíveis, enquanto reinicializações reais do runtime continuarem recuperáveis, sem alterações para OOM e saídas comuns.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

area:compute area:sandbox state:stale

User Story

As an OpenShell operator, I want sandbox status to distinguish an intentional shutdown from a runtime interruption, so that stopped sandboxes are not presented as having restarted unexpectedly and real runtime restarts remain recoverable.

Problem Statement

The Docker and Podman drivers currently classify exits 137 (SIGKILL) and 143 (SIGTERM) as ContainerRuntimeRestart. Those codes establish only that a process was terminated by a signal; they do not identify the sender or intent. An explicit gateway stop that forwards SIGTERM therefore produces the same condition as a Podman/Docker machine or daemon restart.

The durable Stopping phase now prevents that ambiguity from promoting an in-flight explicit stop to Error, but a delayed watcher snapshot can still arrive after Stopped is persisted and replace the user-visible status reason with ContainerRuntimeRestart.

Impact / Why This Matters

Operators can see a sandbox in Stopped phase with a contradictory runtime-restart condition after a normal stop. More broadly, treating all 137/143 exits as runtime restarts conflates graceful stop, forced timeout kill, external intervention, and genuine runtime interruption. The current workaround is to infer intent from lifecycle phase, which protects the immediate flow but does not make the driver status semantically precise.

Acceptance Criteria

  • An explicit gateway stop remains Stopped when a late Docker or Podman signal-exit snapshot arrives, and its terminal status continues to report the intentional stop.
  • A signal termination without explicit stop intent remains distinguishable from a confirmed runtime interruption.
  • Gateway restart recovery continues to recover sandboxes interrupted by a real Docker or Podman runtime/machine restart.
  • OOM termination and ordinary application exits keep their existing distinct behavior.
  • Regression coverage covers Docker and Podman for explicit SIGTERM stop, forced SIGKILL timeout, delayed watcher delivery, OOM, and runtime/machine restart.

Reproduction Steps

  1. Start a Docker- or Podman-backed sandbox.
  2. Stop it through the gateway so the supervisor forwards SIGTERM to its workload.
  3. Observe the driver report exit 143 as ContainerRuntimeRestart.
  4. Deliver that watcher snapshot after the gateway has persisted Stopped.
  5. Observe the sandbox phase remain Stopped while its condition reason no longer reflects the intentional stop.

Environment

  • OpenShell: current main development build
  • Compute drivers: Docker and rootless Podman
  • Related issue: #2855
  • Historical recovery behavior: #2179

Agent Investigation

ContainerRuntimeRestart is currently a heuristic for exit 137/143 in both Docker and Podman. The exit status has no provenance, so operation intent and independently observed runtime state must be considered separately.

Linguagem predominante
Rust
Estrelas
15.4k
Forks
1.7k
Merge médio
1d 21h
PRs com merge (30d)
366

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de NVIDIA/OpenShell

Todas as issues de NVIDIA/OpenShell

Issues semelhantes

Mais issues de Rust

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.