Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

A turn that lost its orchestration lock keeps running, and the same runtime can start a second replay of the instance

オープン
#59 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
38/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
rust

調査の方向性

Start in src/runtime/dispatchers/orchestration.rs, reading spawn_orchestration_lock_renewal_task and the dispatcher loop. Review the regression scenarios described in the SDK PRs, then determine how renewal loss reaches the turn and how in-flight instances are tracked. Done means failed renewals are handled, lost-lock turns do not commit or emit the specified side effects, duplicate replays are avoided, and the loss is logged at warn level with the instance ID.

索引モデルが issue の本文から書いたものです。

説明

bug

Summary

When the lock of a running orchestration turn expires, nothing stops the turn. It runs to its end and only then finds out, because the provider rejects its commit. In the meantime another dispatcher slot can fetch the instance and replay it. That slot can be in the same process. So two replays of one instance can be alive in one process at the same time.

The provider keeps the stored data safe: the commit of the replay without the lock is rejected. But the overlap is a trap for any code that keeps per-instance state in the process.

The Node.js and Python SDKs had such a bug (microsoft/duroxide-node#17, microsoft/duroxide-python#17). A context map keyed by instance ID let the old replay write KV changes into the context of the new replay. The new replay held a valid lock and committed them. The next replay failed with nondeterministic: kv set mismatch.

Where

How the lock gets lost

  • The process stalls for longer than orchestrator_lock_timeout (5 s by default). Causes are memory pressure, a long pause of the process, CPU starvation, or a suspended VM.
  • One renewal call fails or is slow. With a 5 s lock the renewal runs every 3 s, and the renewal task stops at the first error. One miss is enough.

How this was checked

The regression tests in the two SDK PRs reproduce the overlap. The test process stops itself (SIGSTOP) for 7 s in the middle of a turn, with 2 dispatcher slots and the default 5 s lock. After it continues, the old replay and a new replay of the same instance run at the same time in one process.

Suggested fix

  1. Retry a failed renewal while the lock can still be valid. Do not stop at the first retryable error.
  2. Tell the turn when its lock is lost. At least skip the commit, its logs and its metrics. Better: cancel the turn.
  3. Keep a set of the instances that are in flight in this runtime. A slot that fetches an instance from that set abandons the item with a short delay and does not start a second replay.
  4. Log a lost lock at warn level, with the instance ID.

Tracked in #55.

主要言語
Rust
スター
221
フォーク
61
平均マージ
3日 5時間
マージ済み PR(30日)
1

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

microsoft/duroxide のほかの issue

microsoft/duroxide の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。