bug: distinguish intentional signal stops from runtime restarts
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 45/100
- Tipo de issue
- Bug
- Clareza
- Razoavelmente clara
- Status de atividade
- Ativa
- Stack de tecnologia
- docker, rust
- Domínio
- backend, infrastructure
Direção de pesquisa
Rastreie o tratamento dos códigos de saída 137/143 pelos drivers Docker e Podman e, em seguida, acompanhe como a intenção de parada do gateway e os snapshots atrasados do watcher atualizam o status do sandbox. Revise o comportamento de recuperação existente relacionado aos issues #2855 e #2179 e execute ou amplie a cobertura de regressão para ambos os drivers. Considera-se concluído quando as paradas intencionais continuarem sendo terminais e distinguíveis, enquanto reinicializações reais do runtime continuarem recuperáveis, sem alterações para OOM e saídas comuns.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
User Story
As an OpenShell operator, I want sandbox status to distinguish an intentional shutdown from a runtime interruption, so that stopped sandboxes are not presented as having restarted unexpectedly and real runtime restarts remain recoverable.
Problem Statement
The Docker and Podman drivers currently classify exits 137 (SIGKILL) and 143 (SIGTERM) as ContainerRuntimeRestart. Those codes establish only that a process was terminated by a signal; they do not identify the sender or intent. An explicit gateway stop that forwards SIGTERM therefore produces the same condition as a Podman/Docker machine or daemon restart.
The durable Stopping phase now prevents that ambiguity from promoting an in-flight explicit stop to Error, but a delayed watcher snapshot can still arrive after Stopped is persisted and replace the user-visible status reason with ContainerRuntimeRestart.
Impact / Why This Matters
Operators can see a sandbox in Stopped phase with a contradictory runtime-restart condition after a normal stop. More broadly, treating all 137/143 exits as runtime restarts conflates graceful stop, forced timeout kill, external intervention, and genuine runtime interruption. The current workaround is to infer intent from lifecycle phase, which protects the immediate flow but does not make the driver status semantically precise.
Acceptance Criteria
- An explicit gateway stop remains
Stoppedwhen a late Docker or Podman signal-exit snapshot arrives, and its terminal status continues to report the intentional stop. - A signal termination without explicit stop intent remains distinguishable from a confirmed runtime interruption.
- Gateway restart recovery continues to recover sandboxes interrupted by a real Docker or Podman runtime/machine restart.
- OOM termination and ordinary application exits keep their existing distinct behavior.
- Regression coverage covers Docker and Podman for explicit SIGTERM stop, forced SIGKILL timeout, delayed watcher delivery, OOM, and runtime/machine restart.
Reproduction Steps
- Start a Docker- or Podman-backed sandbox.
- Stop it through the gateway so the supervisor forwards SIGTERM to its workload.
- Observe the driver report exit 143 as
ContainerRuntimeRestart. - Deliver that watcher snapshot after the gateway has persisted
Stopped. - Observe the sandbox phase remain
Stoppedwhile its condition reason no longer reflects the intentional stop.
Environment
- OpenShell: current main development build
- Compute drivers: Docker and rootless Podman
- Related issue: #2855
- Historical recovery behavior: #2179
Agent Investigation
ContainerRuntimeRestart is currently a heuristic for exit 137/143 in both Docker and Podman. The exit status has no provenance, so operation intent and independently observed runtime state must be considered separately.
- Linguagem predominante
- Rust
- Estrelas
- 15.4k
- Forks
- 1.7k
- Merge médio
- 1d 21h
- PRs com merge (30d)
- 366
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de NVIDIA/OpenShell
-
state:triage-needed
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
Mantenedores costumam responder em até 1 dia
-
docs: document workspace and provider label capabilitiesTalvez já em andamento @johntmyers assumiu há 3 dias. Abertaarea:docs
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
NVIDIA/OpenShell#4250 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
bug(driver-mxc): test helper fails to compile after gateway-name argumentTalvez já em andamento @feloy assumiu há 4 dias. Abertastate:triage-needed
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
Mantenedores costumam responder em até 1 dia
-
bug: install.sh ignores XDG_CONFIG_HOME for the local gateway configTalvez já em andamento @fede-kamel assumiu há 8 dias. Abertaarea:cli os:linux os:macos state:validated
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
NVIDIA/OpenShell#4042 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
state:triage-needed
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
NVIDIA/OpenShell#3995 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
Todas as issues de NVIDIA/OpenShell
Issues semelhantes
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 62/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 66/100
Mantenedores costumam responder em até 5 dias
-
✨ enhancement needs-discussion
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100
-
virtio-fs (Linux passthrough): debug log in do_lookup panics the fs worker on non-UTF-8 file namesTalvez já em andamento @zcl-g5 assumiu hoje. Aberta
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100
Mantenedores costumam responder em até 2 dias
-
docs(openclaw): RTK_REWRITE_HOST relaxes every default ask, not only commands no rule matchedAbertaarea:docs documentation good first issue priority:low
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
rtk-ai/rtk#4500 · 1 comentário ·
Mantenedores costumam responder em até 1 dia