A failed commit becomes a permanent OrchestrationFailed without parent notification; the fallback abandon never runs
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- rust
- Domain
- distributed-systems
Research direction
Start with process_orchestration_item and ack_orchestration_with_changes in src/runtime/dispatchers/orchestration.rs, then read the related design in #55. Trace the retry and error branches, including both abandon_orchestration_item calls; done means the chosen failure behavior executes asynchronously and does not leave parent notification or activity cancellation missing.
Written by the indexing model from the issue text.
Description
Summary
When ack_orchestration_item fails, the runtime commits an OrchestrationFailed event for the instance. It does this for every error that the provider calls non-retryable, and for a retryable error that lasts longer than about 310 ms. So a database problem that goes away can end an orchestration for good.
The same code path has two more problems:
- The failure is committed without messages. A parent orchestration gets no
SubOrchFailedand waits forever. - The fallback call to
abandon_orchestration_itemnever runs.
Where
process_orchestration_item, the ack and its error branch: https://github.com/microsoft/duroxide/blob/6a458861763a7aa5b78a7c1c97691a6f00489a8b/src/runtime/dispatchers/orchestration.rs#L949-L1017ack_orchestration_with_changes, the retry loop: https://github.com/microsoft/duroxide/blob/6a458861763a7aa5b78a7c1c97691a6f00489a8b/src/runtime/dispatchers/orchestration.rs#L1161-L1210
Details
- Retry budget. A retryable error is retried 5 times, with waits of 10, 20, 40, 80 and 160 ms. Then the error is returned (L1191-L1206). A non-retryable error is returned at once (L1183-L1188).
- Failure commit. The caller builds
OrchestrationFailedfrom the history it fetched and commits it with the same lock token (L975-L1001). If the database is healthy again at that moment, the commit works and the instance isFailed. - No messages. That commit passes empty
worker_items,orchestrator_itemsandcancelled_activities(L990-L999). If the instance is a sub-orchestration, its parent is not told. In-flight activities are not cancelled. - The abandon never runs. Both fallback paths call
drop(self.history_store.abandon_orchestration_item(...))(L1009-L1013 and L1200-L1204).abandon_orchestration_itemis anasync fn. A future that is dropped without.awaitdoes nothing. The lock is only released when it times out.
Which errors are non-retryable depends on the provider. duroxide-pg classifies almost every SQLSTATE as permanent, including a statement timeout: microsoft/duroxide-pg#30.
How this was checked
Found by reading the code. Not reproduced.
Suggested fix
- Do not fail the instance for an infrastructure error on commit. Abandon the item with a backoff, and let
max_attemptsend a message that can never commit. - If the failure commit stays, queue
SubOrchFailedfor the parent and cancel the in-flight activities, as the normal failure path does. .awaitthe twoabandon_orchestration_itemcalls.
Tracked in #55.
- Dominant language
- Rust
- Stars
- 221
- Forks
- 61
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 1
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/duroxide
-
One failed session lock renewal loses the session: no retry, no log, no signal to running activitiesOpenbug
Difficulty 4/5 3-5 days Newbie friendliness 30/100
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 72/100
-
bug
Difficulty 5/5 Over a week Newbie friendliness 38/100
-
prune_kv_values_updated_before emits actions in HashMap order; replay fails with `kv clear mismatch`Possibly taken @akhil9tiet claimed this 3 days ago. Openbug
Difficulty 3/5 1-2 days Newbie friendliness 25/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 35/100
All issues in microsoft/duroxide
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
rubys/roundhouse#571 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 90/100
fastrevmd-lab/rustmistmcp#161 ·
-
arch-audit refactor
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
SocketDev/socket-patch#1011 ·
Maintainers usually reply within 1 day