Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[WORKFLOW SDK FEATURE REQUEST] Retry WaitForInstanceCompletion/Start on a server-sent CANCELLED

未关闭
#1,246 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 3 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
55/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
活跃
技术栈
python

调研方向

Start in dapr/ext/workflow/_durabletask/client.py around lines 335-349 and the matching retry logic in aio/client.py. Read the three cancellation tests in tests/ext/workflow/durabletask/test_orchestration_wait.py and test_client_async.py, then confirm the intended retry and timeout behavior with maintainers. Done means both clients re-issue eligible waits after server CANCELLED responses while preserving caller deadlines and the updated tests pass.

由索引模型根据 Issue 内容生成。

描述

dapr-ext-workflow kind/enhancement

Proposal

wait_for_workflow_completion and wait_for_workflow_start (sync and aio) should re-issue WaitForInstanceCompletion / WaitForInstanceStart when the server ends the call with CANCELLED and the caller's own timeout hasn't expired. Today that status reaches the caller as an error, even though the workflow is still running. This asks maintainers to decide on the behaviour first, because it reverses assertions added in #1112. It's a proposal, not a PR.

Why the server sends CANCELLED

When daprd shuts down or restarts mid-wait (rollout, pod restart, hot-reload, a config change that restarts the runtime), the actors router cancels every in-flight call it tracks. The wait then returns CANCELLED "context canceled", even though the caller is still waiting. The mechanism and a repro are in dapr/dapr#10566.

Reproduced against native daprd 1.18.0:

  • a 60 s workflow with wait_for_workflow_completion(id) and no timeout returns normally;
  • the same run with daprd sent SIGTERM about 10 s in fails with StatusCode.CANCELLED "context canceled".

Why retrying is safe

  • A blocking unary call only gets CANCELLED from the server. The sync client can't cancel it mid-flight, and in the aio client a caller's cancellation raises asyncio.CancelledError, not AioRpcError. A client-side deadline shows up as DEADLINE_EXCEEDED, which already maps to TimeoutError.
  • The wait only reads state, so re-issuing it is idempotent. It returns straight away if the workflow already finished.

What would change

Relationship to the runtime fix

The proper fix is in daprd: return UNAVAILABLE in this case (dapr/dapr#10566). The SDK already retries UNAVAILABLE. This change would protect users on current and older runtimes until that ships, and it would still help afterwards with proxies that reset the stream.

Question for maintainers

Is it acceptable to treat a server-sent CANCELLED on these two waits as retryable, and change the #1112 tests? If yes, it's a small change in both clients plus tests.

主要语言
Python
星标
272
派生
152
平均合并
3 天 13 小时
30 天内合并 PR
11

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

dapr/python-sdk 的其他 Issue

查看 dapr/python-sdk 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。