Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

A turn that lost its orchestration lock keeps running, and the same runtime can start a second replay of the instance

Đang mở
#59 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
38/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
rust

Hướng nghiên cứu

Start in src/runtime/dispatchers/orchestration.rs, reading spawn_orchestration_lock_renewal_task and the dispatcher loop. Review the regression scenarios described in the SDK PRs, then determine how renewal loss reaches the turn and how in-flight instances are tracked. Done means failed renewals are handled, lost-lock turns do not commit or emit the specified side effects, duplicate replays are avoided, and the loss is logged at warn level with the instance ID.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

bug

Summary

When the lock of a running orchestration turn expires, nothing stops the turn. It runs to its end and only then finds out, because the provider rejects its commit. In the meantime another dispatcher slot can fetch the instance and replay it. That slot can be in the same process. So two replays of one instance can be alive in one process at the same time.

The provider keeps the stored data safe: the commit of the replay without the lock is rejected. But the overlap is a trap for any code that keeps per-instance state in the process.

The Node.js and Python SDKs had such a bug (microsoft/duroxide-node#17, microsoft/duroxide-python#17). A context map keyed by instance ID let the old replay write KV changes into the context of the new replay. The new replay held a valid lock and committed them. The next replay failed with nondeterministic: kv set mismatch.

Where

How the lock gets lost

  • The process stalls for longer than orchestrator_lock_timeout (5 s by default). Causes are memory pressure, a long pause of the process, CPU starvation, or a suspended VM.
  • One renewal call fails or is slow. With a 5 s lock the renewal runs every 3 s, and the renewal task stops at the first error. One miss is enough.

How this was checked

The regression tests in the two SDK PRs reproduce the overlap. The test process stops itself (SIGSTOP) for 7 s in the middle of a turn, with 2 dispatcher slots and the default 5 s lock. After it continues, the old replay and a new replay of the same instance run at the same time in one process.

Suggested fix

  1. Retry a failed renewal while the lock can still be valid. Do not stop at the first retryable error.
  2. Tell the turn when its lock is lost. At least skip the commit, its logs and its metrics. Better: cancel the turn.
  3. Keep a set of the instances that are in flight in this runtime. A slot that fetches an instance from that set abandons the item with a short delay and does not start a second replay.
  4. Log a lost lock at warn level, with the instance ID.

Tracked in #55.

Ngôn ngữ chính
Rust
Star
221
Fork
61
Merge trung bình
3 ngày 5 giờ
Pull request đã merge (30 ngày)
1

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của microsoft/duroxide

Tất cả issue của microsoft/duroxide

Issue tương tự

Thêm issue về Rust

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.