[WORKFLOW SDK FEATURE REQUEST] Retry WaitForInstanceCompletion/Start on a server-sent CANCELLED
Los mantenedores suelen responder en 3 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 55/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- backend, distributed-systems
Línea de trabajo
Start in dapr/ext/workflow/_durabletask/client.py around lines 335-349 and the matching retry logic in aio/client.py. Read the three cancellation tests in tests/ext/workflow/durabletask/test_orchestration_wait.py and test_client_async.py, then confirm the intended retry and timeout behavior with maintainers. Done means both clients re-issue eligible waits after server CANCELLED responses while preserving caller deadlines and the updated tests pass.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Proposal
wait_for_workflow_completion and wait_for_workflow_start (sync and aio) should re-issue WaitForInstanceCompletion / WaitForInstanceStart when the server ends the call with CANCELLED and the caller's own timeout hasn't expired. Today that status reaches the caller as an error, even though the workflow is still running. This asks maintainers to decide on the behaviour first, because it reverses assertions added in #1112. It's a proposal, not a PR.
Why the server sends CANCELLED
When daprd shuts down or restarts mid-wait (rollout, pod restart, hot-reload, a config change that restarts the runtime), the actors router cancels every in-flight call it tracks. The wait then returns CANCELLED "context canceled", even though the caller is still waiting. The mechanism and a repro are in dapr/dapr#10566.
Reproduced against native daprd 1.18.0:
- a 60 s workflow with
wait_for_workflow_completion(id)and no timeout returns normally; - the same run with daprd sent
SIGTERMabout 10 s in fails withStatusCode.CANCELLED"context canceled".
Why retrying is safe
- A blocking unary call only gets
CANCELLEDfrom the server. The sync client can't cancel it mid-flight, and in the aio client a caller's cancellation raisesasyncio.CancelledError, notAioRpcError. A client-side deadline shows up asDEADLINE_EXCEEDED, which already maps toTimeoutError. - The wait only reads state, so re-issuing it is idempotent. It returns straight away if the workflow already finished.
What would change
- Add
CANCELLEDto the retried codes, when the caller's timeout hasn't expired, in_durabletask/client.pyL335 and_durabletask/aio/client.pyL230._is_deadline_cancellation(L349) already turns aCANCELLEDafter the deadline intoTimeoutError. - Restart the 30 s transient-retry window after each successful re-issue, rather than anchoring it on the first transient error. Otherwise a sidecar that restarts more than once during a long wait would still fail it.
- These tests from #1112 currently assert that a
CANCELLEDreaches the caller, and would flip:
Relationship to the runtime fix
The proper fix is in daprd: return UNAVAILABLE in this case (dapr/dapr#10566). The SDK already retries UNAVAILABLE. This change would protect users on current and older runtimes until that ships, and it would still help afterwards with proxies that reset the stream.
Question for maintainers
Is it acceptable to treat a server-sent CANCELLED on these two waits as retryable, and change the #1112 tests? If yes, it's a small change in both clients plus tests.
- Lenguaje dominante
- Python
- Estrellas
- 272
- Forks
- 152
- Merge medio
- 3 d 13 h
- PR fusionados (30 d)
- 11
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de dapr/python-sdk
-
dapr-ext-workflow good first issue kind/enhancement P2
Dificultad 2/5 1-2 días Aptitud para principiantes 72/100
dapr/python-sdk#853 · 4 comentarios ·
Los mantenedores suelen responder en 3 días
-
kind/bug
Dificultad 4/5 3-5 días Aptitud para principiantes 65/100
dapr/python-sdk#1233 ·
Los mantenedores suelen responder en 3 días
-
kind/bug
Dificultad 4/5 3-5 días Aptitud para principiantes 68/100
dapr/python-sdk#1232 ·
Los mantenedores suelen responder en 3 días
-
kind/bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
dapr/python-sdk#1230 ·
Los mantenedores suelen responder en 3 días
-
dapr-ext-workflow kind/enhancement
Dificultad 3/5 1-2 días Aptitud para principiantes 78/100
dapr/python-sdk#1213 ·
Los mantenedores suelen responder en 3 días
Todos los issues de dapr/python-sdk
Issues similares
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedAbiertoworkflow
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
Los mantenedores suelen responder en 1 día
-
New Submission: TropWATERAbiertometadata submission
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
-
Wrongly named dashboard variableAbiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
canonical/content-cache-operator#163 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[submission]Abiertosubmission
Dificultad 1/5 Menos de una hora Aptitud para principiantes 65/100
leanprover/lean-eval-submissions#1852 ·
Los mantenedores suelen responder en 1 día