A turn that lost its orchestration lock keeps running, and the same runtime can start a second replay of the instance
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 38/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- rust
- Lĩnh vực
- backend, distributed-systems
Hướng nghiên cứu
Start in src/runtime/dispatchers/orchestration.rs, reading spawn_orchestration_lock_renewal_task and the dispatcher loop. Review the regression scenarios described in the SDK PRs, then determine how renewal loss reaches the turn and how in-flight instances are tracked. Done means failed renewals are handled, lost-lock turns do not commit or emit the specified side effects, duplicate replays are avoided, and the loss is logged at warn level with the instance ID.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
When the lock of a running orchestration turn expires, nothing stops the turn. It runs to its end and only then finds out, because the provider rejects its commit. In the meantime another dispatcher slot can fetch the instance and replay it. That slot can be in the same process. So two replays of one instance can be alive in one process at the same time.
The provider keeps the stored data safe: the commit of the replay without the lock is rejected. But the overlap is a trap for any code that keeps per-instance state in the process.
The Node.js and Python SDKs had such a bug (microsoft/duroxide-node#17, microsoft/duroxide-python#17). A context map keyed by instance ID let the old replay write KV changes into the context of the new replay. The new replay held a valid lock and committed them. The next replay failed with nondeterministic: kv set mismatch.
Where
spawn_orchestration_lock_renewal_task: on any renewal error the task logs atdebuglevel and stops. The turn is not told. https://github.com/microsoft/duroxide/blob/6a458861763a7aa5b78a7c1c97691a6f00489a8b/src/runtime/dispatchers/orchestration.rs#L319-L328- The dispatcher loop: each slot fetches on its own. Nothing checks whether this runtime still runs a turn for the same instance. https://github.com/microsoft/duroxide/blob/6a458861763a7aa5b78a7c1c97691a6f00489a8b/src/runtime/dispatchers/orchestration.rs#L388-L482
How the lock gets lost
- The process stalls for longer than
orchestrator_lock_timeout(5 s by default). Causes are memory pressure, a long pause of the process, CPU starvation, or a suspended VM. - One renewal call fails or is slow. With a 5 s lock the renewal runs every 3 s, and the renewal task stops at the first error. One miss is enough.
How this was checked
The regression tests in the two SDK PRs reproduce the overlap. The test process stops itself (SIGSTOP) for 7 s in the middle of a turn, with 2 dispatcher slots and the default 5 s lock. After it continues, the old replay and a new replay of the same instance run at the same time in one process.
Suggested fix
- Retry a failed renewal while the lock can still be valid. Do not stop at the first retryable error.
- Tell the turn when its lock is lost. At least skip the commit, its logs and its metrics. Better: cancel the turn.
- Keep a set of the instances that are in flight in this runtime. A slot that fetches an instance from that set abandons the item with a short delay and does not start a second replay.
- Log a lost lock at
warnlevel, with the instance ID.
Tracked in #55.
- Ngôn ngữ chính
- Rust
- Star
- 221
- Fork
- 61
- Merge trung bình
- 3 ngày 5 giờ
- Pull request đã merge (30 ngày)
- 1
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/duroxide
-
One failed session lock renewal loses the session: no retry, no log, no signal to running activitiesĐang mởbug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 30/100
-
bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 72/100
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
-
prune_kv_values_updated_before emits actions in HashMap order; replay fails with `kv clear mismatch`Có thể đã có người làm @akhil9tiet đã nhận 3 ngày trước. Đang mởbug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 25/100
-
A failed activity lock renewal can lose the activity result; the orchestration waits foreverĐang mởbug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
Tất cả issue của microsoft/duroxide
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 4 ngày
-
Update dusk-bls12_381 to 0.16Có thể đã có người làm @HDauven đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
googlefonts/fontquant#43 ·
-
bot:ai-assisted status:untriaged
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
midnightntwrk/midnight-zk#561 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
rubys/roundhouse#571 ·
Maintainer thường phản hồi trong vòng 1 ngày