Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

A failed commit becomes a permanent OrchestrationFailed without parent notification; the fallback abandon never runs

Đang mở
#58 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
48/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
rust
Lĩnh vực
distributed-systems

Hướng nghiên cứu

Start with process_orchestration_item and ack_orchestration_with_changes in src/runtime/dispatchers/orchestration.rs, then read the related design in #55. Trace the retry and error branches, including both abandon_orchestration_item calls; done means the chosen failure behavior executes asynchronously and does not leave parent notification or activity cancellation missing.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

bug

Summary

When ack_orchestration_item fails, the runtime commits an OrchestrationFailed event for the instance. It does this for every error that the provider calls non-retryable, and for a retryable error that lasts longer than about 310 ms. So a database problem that goes away can end an orchestration for good.

The same code path has two more problems:

  • The failure is committed without messages. A parent orchestration gets no SubOrchFailed and waits forever.
  • The fallback call to abandon_orchestration_item never runs.

Where

Details

  1. Retry budget. A retryable error is retried 5 times, with waits of 10, 20, 40, 80 and 160 ms. Then the error is returned (L1191-L1206). A non-retryable error is returned at once (L1183-L1188).
  2. Failure commit. The caller builds OrchestrationFailed from the history it fetched and commits it with the same lock token (L975-L1001). If the database is healthy again at that moment, the commit works and the instance is Failed.
  3. No messages. That commit passes empty worker_items, orchestrator_items and cancelled_activities (L990-L999). If the instance is a sub-orchestration, its parent is not told. In-flight activities are not cancelled.
  4. The abandon never runs. Both fallback paths call drop(self.history_store.abandon_orchestration_item(...)) (L1009-L1013 and L1200-L1204). abandon_orchestration_item is an async fn. A future that is dropped without .await does nothing. The lock is only released when it times out.

Which errors are non-retryable depends on the provider. duroxide-pg classifies almost every SQLSTATE as permanent, including a statement timeout: microsoft/duroxide-pg#30.

How this was checked

Found by reading the code. Not reproduced.

Suggested fix

  • Do not fail the instance for an infrastructure error on commit. Abandon the item with a backoff, and let max_attempts end a message that can never commit.
  • If the failure commit stays, queue SubOrchFailed for the parent and cancel the in-flight activities, as the normal failure path does.
  • .await the two abandon_orchestration_item calls.

Tracked in #55.

Ngôn ngữ chính
Rust
Star
221
Fork
61
Merge trung bình
3 ngày 5 giờ
Pull request đã merge (30 ngày)
1

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của microsoft/duroxide

Tất cả issue của microsoft/duroxide

Issue tương tự

Thêm issue về Rust

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.