bug: distinguish intentional signal stops from runtime restarts
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- docker, rust
- Lĩnh vực
- backend, infrastructure
Hướng nghiên cứu
Theo dõi cách các driver Docker và Podman xử lý các mã thoát 137/143, sau đó theo dõi cách ý định dừng của gateway và các snapshot watcher bị trì hoãn cập nhật trạng thái sandbox. Xem xét hành vi khôi phục hiện có liên quan đến các issue #2855 và #2179, đồng thời chạy hoặc mở rộng phạm vi kiểm thử hồi quy cho cả hai driver. Hoàn thành có nghĩa là các lần dừng có chủ đích vẫn ở trạng thái kết thúc và có thể phân biệt được, trong khi các lần khởi động lại runtime thực sự vẫn có thể khôi phục, còn OOM và các lần thoát thông thường không thay đổi.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
User Story
As an OpenShell operator, I want sandbox status to distinguish an intentional shutdown from a runtime interruption, so that stopped sandboxes are not presented as having restarted unexpectedly and real runtime restarts remain recoverable.
Problem Statement
The Docker and Podman drivers currently classify exits 137 (SIGKILL) and 143 (SIGTERM) as ContainerRuntimeRestart. Those codes establish only that a process was terminated by a signal; they do not identify the sender or intent. An explicit gateway stop that forwards SIGTERM therefore produces the same condition as a Podman/Docker machine or daemon restart.
The durable Stopping phase now prevents that ambiguity from promoting an in-flight explicit stop to Error, but a delayed watcher snapshot can still arrive after Stopped is persisted and replace the user-visible status reason with ContainerRuntimeRestart.
Impact / Why This Matters
Operators can see a sandbox in Stopped phase with a contradictory runtime-restart condition after a normal stop. More broadly, treating all 137/143 exits as runtime restarts conflates graceful stop, forced timeout kill, external intervention, and genuine runtime interruption. The current workaround is to infer intent from lifecycle phase, which protects the immediate flow but does not make the driver status semantically precise.
Acceptance Criteria
- An explicit gateway stop remains
Stoppedwhen a late Docker or Podman signal-exit snapshot arrives, and its terminal status continues to report the intentional stop. - A signal termination without explicit stop intent remains distinguishable from a confirmed runtime interruption.
- Gateway restart recovery continues to recover sandboxes interrupted by a real Docker or Podman runtime/machine restart.
- OOM termination and ordinary application exits keep their existing distinct behavior.
- Regression coverage covers Docker and Podman for explicit SIGTERM stop, forced SIGKILL timeout, delayed watcher delivery, OOM, and runtime/machine restart.
Reproduction Steps
- Start a Docker- or Podman-backed sandbox.
- Stop it through the gateway so the supervisor forwards SIGTERM to its workload.
- Observe the driver report exit 143 as
ContainerRuntimeRestart. - Deliver that watcher snapshot after the gateway has persisted
Stopped. - Observe the sandbox phase remain
Stoppedwhile its condition reason no longer reflects the intentional stop.
Environment
- OpenShell: current main development build
- Compute drivers: Docker and rootless Podman
- Related issue: #2855
- Historical recovery behavior: #2179
Agent Investigation
ContainerRuntimeRestart is currently a heuristic for exit 137/143 in both Docker and Podman. The exit status has no provenance, so operation intent and independently observed runtime state must be considered separately.
- Ngôn ngữ chính
- Rust
- Star
- 8.7k
- Fork
- 1.3k
- Merge trung bình
- 2 ngày 6 giờ
- Pull request đã merge (30 ngày)
- 297
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/OpenShell
-
area:docs
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
-
state:triage-needed
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
area:cli state:validated
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
state:triage-needed
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
-
area:build spike state:review-ready state:stale
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 68/100
Tất cả issue của NVIDIA/OpenShell
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
state:needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
zed-industries/zed#64680 · 2 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
RustPython/RustPython#8802 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
TheLarkInn/aipm#2390 ·