Stale activity cleanup: TTL or background sweep for undeliverable worker queue items
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- postgresql, rust
- Domain
- backend, databases, distributed-systems
Research direction
Start with the TODO.md entry and trace the provider trait, worker queue items, RuntimeOptions, and activity builder described in the issue, including the CosmosDB and Postgres providers. Decide which expiry design and retry semantics fit the existing error infrastructure, then verify that expired items are failed back to the orchestrator while the default remains no expiry.
Written by the indexing model from the issue text.
Description
Problem
Activities scheduled with a tag that no worker is configured to handle (or activities whose target worker goes offline permanently) will sit in the worker queue indefinitely. The orchestration that scheduled them hangs forever unless the user manually implements a select2(activity, timer) starvation guard.
This is especially relevant now that activity tagging is implemented -- it is easy to schedule a .with_tag("gpu") activity in an environment where no GPU worker is running.
Desired Behavior
Undeliverable or stale activities should not block orchestrations forever. The runtime should detect activities that exceed a configurable time limit and fail them back to the orchestrator with a clear error.
Proposed Approaches
Option A: Activity TTL (per-item expiry)
- Add an optional
expires_attimestamp to worker queue items (set at enqueue time based on a configurable TTL) - Provider
fetch_work_item()skips expired items - A periodic sweep (or check at fetch time) marks expired items as failed
- The orchestration receives an
ActivityExpirederror it can match on
Option B: Background cleanup process
- A runtime background task periodically scans for worker queue items older than a configurable threshold
- Stale items are failed back to the orchestrator with a timeout error
- Simpler to implement but less granular (global threshold vs per-activity)
Option C: Hybrid
- Default global TTL from
RuntimeOptions(e.g., 1 hour) - Per-activity override via
.with_ttl(Duration)on the activity builder - Background sweep handles the cleanup
Design Considerations
- Provider trait changes: Need
expires_atfield or equivalent on worker queue items - Event model: New
ActivityExpiredor reuse existing error infrastructure - Backward compatibility: TTL should be optional, default to no expiry (current behavior) for existing users
- CosmosDB / Postgres providers: Both need the expiry field; CosmosDB has native TTL support that could be leveraged
- Interaction with retries: Should TTL apply per-attempt or total? Probably total elapsed since first enqueue
Related
- Activity tagging feature (
.with_tag()/TagFilter) select2(activity, timer)starvation-safe pattern (current workaround)- TODO.md entry added for tracking
- Dominant language
- Rust
- Stars
- 217
- Forks
- 61
- Avg merge
- 3d 9m
- Merged PRs (30d)
- 2
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/duroxide
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 68/100
All issues in microsoft/duroxide
Similar issues
-
Change output crossing a compactsize boundary leaves the fee slightly below the requested feerateOpenbug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
bitcoindevkit/bdk_wallet#578 ·
Maintainers usually reply within 8 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
521xueweihan/HelloGitHub#3832 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
canonical/opentelemetry-collector-operator#409 ·
Maintainers usually reply within 1 day