Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

A failed commit becomes a permanent OrchestrationFailed without parent notification; the fallback abandon never runs

Open
#58 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
rust

Research direction

Start with process_orchestration_item and ack_orchestration_with_changes in src/runtime/dispatchers/orchestration.rs, then read the related design in #55. Trace the retry and error branches, including both abandon_orchestration_item calls; done means the chosen failure behavior executes asynchronously and does not leave parent notification or activity cancellation missing.

Written by the indexing model from the issue text.

Description

bug

Summary

When ack_orchestration_item fails, the runtime commits an OrchestrationFailed event for the instance. It does this for every error that the provider calls non-retryable, and for a retryable error that lasts longer than about 310 ms. So a database problem that goes away can end an orchestration for good.

The same code path has two more problems:

  • The failure is committed without messages. A parent orchestration gets no SubOrchFailed and waits forever.
  • The fallback call to abandon_orchestration_item never runs.

Where

Details

  1. Retry budget. A retryable error is retried 5 times, with waits of 10, 20, 40, 80 and 160 ms. Then the error is returned (L1191-L1206). A non-retryable error is returned at once (L1183-L1188).
  2. Failure commit. The caller builds OrchestrationFailed from the history it fetched and commits it with the same lock token (L975-L1001). If the database is healthy again at that moment, the commit works and the instance is Failed.
  3. No messages. That commit passes empty worker_items, orchestrator_items and cancelled_activities (L990-L999). If the instance is a sub-orchestration, its parent is not told. In-flight activities are not cancelled.
  4. The abandon never runs. Both fallback paths call drop(self.history_store.abandon_orchestration_item(...)) (L1009-L1013 and L1200-L1204). abandon_orchestration_item is an async fn. A future that is dropped without .await does nothing. The lock is only released when it times out.

Which errors are non-retryable depends on the provider. duroxide-pg classifies almost every SQLSTATE as permanent, including a statement timeout: microsoft/duroxide-pg#30.

How this was checked

Found by reading the code. Not reproduced.

Suggested fix

  • Do not fail the instance for an infrastructure error on commit. Abandon the item with a backoff, and let max_attempts end a message that can never commit.
  • If the failure commit stays, queue SubOrchFailed for the parent and cancel the in-flight activities, as the normal failure path does.
  • .await the two abandon_orchestration_item calls.

Tracked in #55.

Dominant language
Rust
Stars
221
Forks
61
Avg merge
3d 5h
Merged PRs (30d)
1

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/duroxide

All issues in microsoft/duroxide

Similar issues

More Rust issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.