A turn that lost its orchestration lock keeps running, and the same runtime can start a second replay of the instance
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 38/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- rust
- Domain
- backend, distributed-systems
Research direction
Start in src/runtime/dispatchers/orchestration.rs, reading spawn_orchestration_lock_renewal_task and the dispatcher loop. Review the regression scenarios described in the SDK PRs, then determine how renewal loss reaches the turn and how in-flight instances are tracked. Done means failed renewals are handled, lost-lock turns do not commit or emit the specified side effects, duplicate replays are avoided, and the loss is logged at warn level with the instance ID.
Written by the indexing model from the issue text.
Description
Summary
When the lock of a running orchestration turn expires, nothing stops the turn. It runs to its end and only then finds out, because the provider rejects its commit. In the meantime another dispatcher slot can fetch the instance and replay it. That slot can be in the same process. So two replays of one instance can be alive in one process at the same time.
The provider keeps the stored data safe: the commit of the replay without the lock is rejected. But the overlap is a trap for any code that keeps per-instance state in the process.
The Node.js and Python SDKs had such a bug (microsoft/duroxide-node#17, microsoft/duroxide-python#17). A context map keyed by instance ID let the old replay write KV changes into the context of the new replay. The new replay held a valid lock and committed them. The next replay failed with nondeterministic: kv set mismatch.
Where
spawn_orchestration_lock_renewal_task: on any renewal error the task logs atdebuglevel and stops. The turn is not told. https://github.com/microsoft/duroxide/blob/6a458861763a7aa5b78a7c1c97691a6f00489a8b/src/runtime/dispatchers/orchestration.rs#L319-L328- The dispatcher loop: each slot fetches on its own. Nothing checks whether this runtime still runs a turn for the same instance. https://github.com/microsoft/duroxide/blob/6a458861763a7aa5b78a7c1c97691a6f00489a8b/src/runtime/dispatchers/orchestration.rs#L388-L482
How the lock gets lost
- The process stalls for longer than
orchestrator_lock_timeout(5 s by default). Causes are memory pressure, a long pause of the process, CPU starvation, or a suspended VM. - One renewal call fails or is slow. With a 5 s lock the renewal runs every 3 s, and the renewal task stops at the first error. One miss is enough.
How this was checked
The regression tests in the two SDK PRs reproduce the overlap. The test process stops itself (SIGSTOP) for 7 s in the middle of a turn, with 2 dispatcher slots and the default 5 s lock. After it continues, the old replay and a new replay of the same instance run at the same time in one process.
Suggested fix
- Retry a failed renewal while the lock can still be valid. Do not stop at the first retryable error.
- Tell the turn when its lock is lost. At least skip the commit, its logs and its metrics. Better: cancel the turn.
- Keep a set of the instances that are in flight in this runtime. A slot that fetches an instance from that set abandons the item with a short delay and does not start a second replay.
- Log a lost lock at
warnlevel, with the instance ID.
Tracked in #55.
- Dominant language
- Rust
- Stars
- 221
- Forks
- 61
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 1
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/duroxide
-
One failed session lock renewal loses the session: no retry, no log, no signal to running activitiesOpenbug
Difficulty 4/5 3-5 days Newbie friendliness 30/100
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 72/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
prune_kv_values_updated_before emits actions in HashMap order; replay fails with `kv clear mismatch`Possibly taken @akhil9tiet claimed this 2 days ago. Openbug
Difficulty 3/5 1-2 days Newbie friendliness 25/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 35/100
All issues in microsoft/duroxide
Similar issues
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
zcashlabs/thus-spoke-zakura#153 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 79/100
topgrade-rs/topgrade#2395 ·
Maintainers usually reply within 1 day
-
app bug windows-os
Difficulty 2/5 1-3 hours Newbie friendliness 67/100
Maintainers usually reply within 1 day
-
editor good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
funnyboy-roks/inq#54 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day