Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Azure Storage backend: control queue partition left unowned for hours/days after lease expires

未关闭
#1,389 1 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 3 天内回复

@nytian 已经在做这个了。

开始于 2026年9月1日。

评估

这个 Issue 还没有评估数据。

描述

Environment
Microsoft.Azure.DurableTask.AzureStorage: 2.9.1
Microsoft.Azure.DurableTask.Core: 3.9.0
Runtime: .NET 8
Host: Kubernetes, self-hosted TaskHubWorker (not Azure Functions)
Partition manager: Table partition manager
PartitionCount: 16
Worker replicas: 25
UseAppLease: default (true)
Non-default settings: TaskHubName, PartitionCount, StorageAccountClientProvider, LoggerFactory only

Summary
A single control queue partition stopped being polled by any worker for hours/days. Messages continued to be enqueued to it, but nothing dequeued them until an unrelated worker eventually acquired the partition, at which point many stranded messages were drained at once (many were  ExecutionStarted). Maximum observed message delay was 3 days for some cases.
No worker process crashed, no node was lost, and the previous owner never released the partition.
For example, sharing one request timeline for one partition:

Timeline (one partition, referred to below as control-NN)

T+00:07 Lease thrash: acquire → LeaseLost → DrainTablePartitionAsync → re-acquire → DropLostControlQueue. Occurs
T+00:21 twice.
T+02:26 Last message consumed from this partition. No lease release is logged.
T+02:26 16 messages enqueued. Zero dequeues.
T+17:23
T+17:23 A worker acquires the partition and drains 17 stranded messages.

Evidence the partition was unowned
No new messages were found backing off is emitted by the control queue reader loop, so its rate tracks fetch attempts ~1:1 whenever a partition is owned

Hour │ Fetches │ Backoff log lines │
01 │ 62 │ 61 │
02 │ 2 │ 2 │
04–15 │ 0 │ 0 │
18 │ 18 │ 18 │

Zero backoff lines during the outage rules out "owned but concurrency-starved" no worker was running a reader loop against this partition at all.
With  LeaseAcquireInterval = 10s and LeaseInterval = 30s, expected reclaim after an owner stops renewing is ~10–30 seconds.

We were expecting to behave the partition normally, but it remained unowned for ~15.3 hours while the other workers continued polling the remaining partitions on the same task hub normally.

Impact
Orchestrations whose  ExecutionStarted  message landed on the affected partition remained in Pending indefinitely, then began executing 10–15 hours later. By that point the external resources they operate on had already been cleaned up by a separate lifecycle process, so the orchestrations failed with misleading downstream errors, which made the underlying cause difficult to attribute.
Activities were unaffected, since the work-item queue is unpartitioned and polled by all workers. Only orchestration-level messages (which ride the partitioned control queues) were impacted.

Recurrence
Occurring regularly for with about

Questions

  1. Is this a known defect in the table partition manager in 2.9.x, and is it addressed in a later release?
  2. Can DropLostControlQueue/DrainTablePartitionAsync  following a  LeaseLost  leave a partition in a state where no worker re-acquires it, or where the owner record isn't cleaned up?
  3. Is running more workers (25) than partitions (16) known to aggravate lease contention in this path? We are considering to reduced replicas to match PartitionCount in our next change.
  4. Is UseAppLease = true appropriate for self-hosted (non-Functions) deployments?
  5. Any recommended mitigation or detection while a fix is pending?

Mitigation / detection we are considering in next change
• Restructured orchestrations to prefer activities over sub-orchestrations, so more work rides the unpartitioned work-item queue
• Reduced worker replicas to match  PartitionCount 
• Added a monitor that alerts when instances remain in Pending beyond a threshold, since a never-dispatched orchestration emits no telemetry of its own

I can provide additional detail if required.

主要语言
C#
星标
1.7k
派生
335
平均合并
5 天 3 小时
30 天内合并 PR
8

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

Azure/durabletask 的其他 Issue

查看 Azure/durabletask 的全部 Issue

相似的 Issue

更多 C# Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。