Azure Storage backend: control queue partition left unowned for hours/days after lease expires
维护者通常 3 天内回复
@nytian 已经在做这个了。
开始于 2026年9月1日。
评估
这个 Issue 还没有评估数据。
描述
Environment
Microsoft.Azure.DurableTask.AzureStorage: 2.9.1
Microsoft.Azure.DurableTask.Core: 3.9.0
Runtime: .NET 8
Host: Kubernetes, self-hosted TaskHubWorker (not Azure Functions)
Partition manager: Table partition manager
PartitionCount: 16
Worker replicas: 25
UseAppLease: default (true)
Non-default settings: TaskHubName, PartitionCount, StorageAccountClientProvider, LoggerFactory only
Summary
A single control queue partition stopped being polled by any worker for hours/days. Messages continued to be enqueued to it, but nothing dequeued them until an unrelated worker eventually acquired the partition, at which point many stranded messages were drained at once (many were ExecutionStarted). Maximum observed message delay was 3 days for some cases.
No worker process crashed, no node was lost, and the previous owner never released the partition.
For example, sharing one request timeline for one partition:
Timeline (one partition, referred to below as control-NN)
T+00:07 Lease thrash: acquire → LeaseLost → DrainTablePartitionAsync → re-acquire → DropLostControlQueue. Occurs
T+00:21 twice.
T+02:26 Last message consumed from this partition. No lease release is logged.
T+02:26 16 messages enqueued. Zero dequeues.
T+17:23
T+17:23 A worker acquires the partition and drains 17 stranded messages.
Evidence the partition was unowned
No new messages were found backing off is emitted by the control queue reader loop, so its rate tracks fetch attempts ~1:1 whenever a partition is owned
Hour │ Fetches │ Backoff log lines │
01 │ 62 │ 61 │
02 │ 2 │ 2 │
04–15 │ 0 │ 0 │
18 │ 18 │ 18 │
Zero backoff lines during the outage rules out "owned but concurrency-starved" no worker was running a reader loop against this partition at all.
With LeaseAcquireInterval = 10s and LeaseInterval = 30s, expected reclaim after an owner stops renewing is ~10–30 seconds.
We were expecting to behave the partition normally, but it remained unowned for ~15.3 hours while the other workers continued polling the remaining partitions on the same task hub normally.
Impact
Orchestrations whose ExecutionStarted message landed on the affected partition remained in Pending indefinitely, then began executing 10–15 hours later. By that point the external resources they operate on had already been cleaned up by a separate lifecycle process, so the orchestrations failed with misleading downstream errors, which made the underlying cause difficult to attribute.
Activities were unaffected, since the work-item queue is unpartitioned and polled by all workers. Only orchestration-level messages (which ride the partitioned control queues) were impacted.
Recurrence
Occurring regularly for with about
Questions
- Is this a known defect in the table partition manager in 2.9.x, and is it addressed in a later release?
- Can DropLostControlQueue/DrainTablePartitionAsync following a LeaseLost leave a partition in a state where no worker re-acquires it, or where the owner record isn't cleaned up?
- Is running more workers (25) than partitions (16) known to aggravate lease contention in this path? We are considering to reduced replicas to match PartitionCount in our next change.
- Is UseAppLease = true appropriate for self-hosted (non-Functions) deployments?
- Any recommended mitigation or detection while a fix is pending?
Mitigation / detection we are considering in next change
• Restructured orchestrations to prefer activities over sub-orchestrations, so more work rides the unpartitioned work-item queue
• Reduced worker replicas to match PartitionCount
• Added a monitor that alerts when instances remain in Pending beyond a threshold, since a never-dispatched orchestration emits no telemetry of its own
I can provide additional detail if required.
- 主要语言
- C#
- 星标
- 1.7k
- 派生
- 335
- 平均合并
- 5 天 3 小时
- 30 天内合并 PR
- 8
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
Azure/durabletask 的其他 Issue
-
难度 4/5 3-5 天 新手友好度 55/100
Azure/durabletask#1398 · 2 条评论 ·
维护者通常 3 天内回复
-
难度 4/5 3-5 天 新手友好度 48/100
Azure/durabletask#1332 ·
维护者通常 3 天内回复
-
难度 3/5 1-2 天 新手友好度 58/100
Azure/durabletask#1318 · 1 条评论 ·
维护者通常 3 天内回复
-
难度 5/5 一周以上 新手友好度 25/100
Azure/durabletask#1301 · 3 条评论 ·
维护者通常 3 天内回复
-
难度 5/5 一周以上 新手友好度 25/100
Azure/durabletask#1297 ·
维护者通常 3 天内回复
查看 Azure/durabletask 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 68/100
stryker-mutator/stryker-net#3892 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
MobiFlight/MobiFlight-Connector#3419 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
Kryptos-FR/MarkView.Avalonia#105 ·
维护者通常 1 天内回复
-
[辞書]未关闭提案 辞書
难度 2/5 1-3 小时 新手友好度 65/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 76/100
microsoft/fluentui-blazor#5410 ·
维护者通常 1 天内回复