[FEATURE] Waiter semantics for `callback`: stop on terminal failure, time-based bounds, backoff
还没有人认领这个 Issue。
评估
调研方向
Start with run_callback_poll in src/core/utils.rs, then trace run_callback and run_troubleshoot in src/commands/base.rs and the callback paths in src/commands/build.rs and src/commands/teardown.rs. Review the documented callback example in website/docs/resource-query-files.md. Done means terminal failures stop polling with diagnostics and troubleshooting, time-based bounds and backoff work, and existing callback behavior remains compatible.
由索引模型根据 Issue 内容生成。
描述
Problem
Asynchronous mutations return a handle that has to be polled until the operation finishes. awscc (Cloud Control) is the main case: every create, update and delete returns a RequestToken. callback is the hook for polling it, but a callback result has only two states: done, or keep polling. There is no way to say "this operation reached a terminal failure, stop". The docs already list this under the callback's known limitations.
Here is a real case. An awscc.s3.buckets create for a bucket name deleted about 40 minutes earlier in another region. Cloud Control accepted the request, reported IN_PROGRESS (with an interim ErrorCode: ServiceInternalError) for about 3 minutes while it retried the S3 call internally, and then finished:
OperationStatus: FAILED
ErrorCode: ResourceConflict
StatusMessage: A conflicting conditional operation is currently in progress against this resource. Please try again. (Service: S3, Status Code: 409, ...)
None of the ways to handle this today reports that message:
- Callback that succeeds on
OperationStatus = 'SUCCESS'(the documented pattern). It keeps polling a request that has alreadyFAILEDuntilretriesruns out, then exits withcallback timeout for [x] create operation after N retries.run_callbackexits throughcatch_error_and_exitwithout runningtroubleshoot, so the reason is never shown. - No callback, relying on the post-create exists check plus
troubleshoot:create. The exists check gives up after the statecheck'sretriesxretry_delay(50 s here).troubleshootthen runs once, while the request is stillIN_PROGRESS. The log showedIN_PROGRESS/ServiceInternalError/StatusMessage: NULL, which isn't the reason. - Workaround: a callback that treats any terminal state as done,
OperationStatus IN ('SUCCESS', 'FAILED', 'CANCEL_COMPLETE') AS success.troubleshootdoes then report the real message. But the post-create exists check still polls for a resource that will never exist (another 50 s), and the final error isnot found after create post-deploy check, create operation may have failed, although the failure was known and explained a minute earlier.
Two more problems with the poll budget:
- It is a count, not a time. The budget is
retriesx a fixedretry_delay. Authors have to guess a count that covers the slowest normal case, which makes failures slow to report and still doesn't guarantee coverage. - Provider hints are ignored. Cloud Control returns
RetryAfterin the progress event, and nothing can use it.
Other providers have the same shape: Azure long-running operations (provisioningState = 'Failed'), Google operations resources (done = true with an error), and Databricks resources with a FAILED / ERROR lifecycle state.
Proposal
Give callback waiter semantics, modelled on AWS SDK waiters, where each poll result is success, failure or retry. Extending the existing hook keeps one polling construct instead of two, and it stays backward compatible.
- Three-state result. The callback query can return an optional
failedcolumn alongsidesuccess:successtruthy -> done (as today)failedtruthy -> terminal failure- neither -> still pending, poll again
- Stop at once on failure. Stop polling and log the whole row as diagnostics, so extra columns such as
messageorerror_codeappear in the output. Then runtroubleshoot:<op>if one is defined, and fail the resource under the existing--on-failurehandling. Skip the post-deploy exists and statecheck for that resource, since the outcome is already known. - Run
troubleshooton callback timeout, before exiting. - Time-based bounds.
timeout=<seconds>: a wall-clock budget, alongside or instead ofretries.backoff=<multiplier>andmax_delay=<seconds>: grow the delay between polls.
- Honour provider hints (optional). If the query returns a
retry_aftercolumn (seconds, or an epoch timestamp, which is what Cloud Control'sRetryAfteris), use it for the next delay, capped bymax_delayandtimeout.
For awscc the whole contract becomes:
/*+ create */
INSERT INTO awscc.s3.buckets (BucketName, region)
SELECT '{{ bucket_name }}', '{{ region }}'
RETURNING *
/*+ callback:create, timeout=600, retry_delay=5, backoff=1.5, max_delay=30 */
SELECT
OperationStatus = 'SUCCESS' AS success,
OperationStatus IN ('FAILED', 'CANCEL_COMPLETE') AS failed,
ErrorCode AS error_code,
StatusMessage AS message
FROM awscc.cloud_control.resource_request
WHERE RequestToken = '{{ callback.RequestToken }}'
AND region = '{{ region }}'
A normal create finishes on the first or second poll. A failed one stops as soon as Cloud Control reports FAILED and prints ResourceConflict plus the S3 message. The run fails on that error, not on a later "not found".
Behaviour to pin down
successandfailedboth truthy. Suggest treating it as a failure (the conservative choice) and logging a warning, since the query is ambiguous.- Query errors and empty results while polling. Treat both as pending, as today: some providers briefly return 404 for a handle just after dispatch. Log the error at
debuglevel on each attempt, and include the last one in the timeout message. retriesandtimeoutboth set. Stop at whichever limit comes first. With neither set, keep the current defaults (retries=3,retry_delay=5).- Short circuit.
short_circuit_field/short_circuit_valuekeep working. The same three-state check could run on theRETURNING *row, so a request that fails synchronously never needs a poll. - Teardown.
callback:deletegets the same semantics. A terminal failure means the resource is reported as not confirmed deleted under--on-failure ignoreand aborts under--on-failure error. - Dry run. Unchanged: callbacks are rendered and logged, not executed.
- Out of scope. A
failedcolumn onstatecheck/exists(for example AzureprovisioningState = 'Failed') could reuse the same evaluation later, but it's a separate change.
Related
website/docs/resource-query-files.md, callback section: the "no mechanism to short-circuit retries on a terminal failure" known limitation goes away.- The same page's
awscccallback example is out of date. It pollsawscc.cloudcontrol.resource_request_statuseswith{{ callback.ProgressEvent.RequestToken }}. The current provider resource isawscc.cloud_control.resource_request, andRETURNING *returns the progress event fields flat, so it's{{ callback.RequestToken }}(theexamples/aws/sqlserverstack uses the flatRequestToken). Worth fixing when this lands, or on its own. - Code:
run_callback_pollinsrc/core/utils.rs,run_callback/run_troubleshootinsrc/commands/base.rs, the create and update callback blocks and the post-deploy exists -> troubleshoot path insrc/commands/build.rs, andcallback:deleteinsrc/commands/teardown.rs.
- 主要语言
- Rust
- 星标
- 1
- 派生
- 0
- 平均合并
- 12 分钟
- 30 天内合并 PR
- 3
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
stackql/stackql-deploy-rs 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 75/100
stackql/stackql-deploy-rs#59 · 1 条评论 ·
-
难度 5/5 一周以上 新手友好度 35/100
stackql/stackql-deploy-rs#58 ·
查看 stackql/stackql-deploy-rs 的全部 Issue
相似的 Issue
-
bug github_actions
难度 2/5 1-3 小时 新手友好度 75/100
registrystack/registry-stack#1393 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
longbridge/gpui-kit#3223 ·
-
bug engine
难度 2/5 1-3 小时 新手友好度 65/100
rocky-data/rocky#2181 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
oasisprotocol/oasis-sdk#2523 ·
-
bot:ai-assisted component:indexer QA-roadmap status:untriaged
难度 2/5 1-3 小时 新手友好度 75/100
midnightntwrk/midnight-indexer#1557 ·