Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

A turn that lost its orchestration lock keeps running, and the same runtime can start a second replay of the instance

Open
#59 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
38/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
rust

Research direction

Start in src/runtime/dispatchers/orchestration.rs, reading spawn_orchestration_lock_renewal_task and the dispatcher loop. Review the regression scenarios described in the SDK PRs, then determine how renewal loss reaches the turn and how in-flight instances are tracked. Done means failed renewals are handled, lost-lock turns do not commit or emit the specified side effects, duplicate replays are avoided, and the loss is logged at warn level with the instance ID.

Written by the indexing model from the issue text.

Description

bug

Summary

When the lock of a running orchestration turn expires, nothing stops the turn. It runs to its end and only then finds out, because the provider rejects its commit. In the meantime another dispatcher slot can fetch the instance and replay it. That slot can be in the same process. So two replays of one instance can be alive in one process at the same time.

The provider keeps the stored data safe: the commit of the replay without the lock is rejected. But the overlap is a trap for any code that keeps per-instance state in the process.

The Node.js and Python SDKs had such a bug (microsoft/duroxide-node#17, microsoft/duroxide-python#17). A context map keyed by instance ID let the old replay write KV changes into the context of the new replay. The new replay held a valid lock and committed them. The next replay failed with nondeterministic: kv set mismatch.

Where

How the lock gets lost

  • The process stalls for longer than orchestrator_lock_timeout (5 s by default). Causes are memory pressure, a long pause of the process, CPU starvation, or a suspended VM.
  • One renewal call fails or is slow. With a 5 s lock the renewal runs every 3 s, and the renewal task stops at the first error. One miss is enough.

How this was checked

The regression tests in the two SDK PRs reproduce the overlap. The test process stops itself (SIGSTOP) for 7 s in the middle of a turn, with 2 dispatcher slots and the default 5 s lock. After it continues, the old replay and a new replay of the same instance run at the same time in one process.

Suggested fix

  1. Retry a failed renewal while the lock can still be valid. Do not stop at the first retryable error.
  2. Tell the turn when its lock is lost. At least skip the commit, its logs and its metrics. Better: cancel the turn.
  3. Keep a set of the instances that are in flight in this runtime. A slot that fetches an instance from that set abandons the item with a short delay and does not start a second replay.
  4. Log a lost lock at warn level, with the instance ID.

Tracked in #55.

Dominant language
Rust
Stars
221
Forks
61
Avg merge
3d 5h
Merged PRs (30d)
1

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/duroxide

All issues in microsoft/duroxide

Similar issues

More Rust issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.