[Bug]: get() on a SKIPPED parallel branch never returns, and the execution fails
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 52/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
- 技术栈
- java
调研方向
Start with ParallelOperation.handleCompletion and ParallelOperation.branch(), then inspect ChildContextOperation.handleChildContextSuccess and the sdk-integration-tests using LocalDurableTestRunner. Compare both skipped-branch cases across the first invocation and replay, and update docs/core/parallel.md as needed. Done means get() consistently returns or throws a named branch exception without producing PENDING with no pending operation.
由索引模型根据 Issue 内容生成。
描述
Expected Behavior
get() on a branch future returns the branch result or throws an exception. It behaves the same way in the first invocation and on every replay.
The parallel result can record a branch as SKIPPED. For such a branch, get() should throw an exception that names the branch. docs/core/parallel.md should document it. For comparison, docs/core/map.md defines a skipped map item: status SKIPPED, a null result, and a null error.
Actual Behavior
get() on a SKIPPED branch does not return, and it does not throw a branch exception. It waits until the SDK suspends the invocation. The invocation then returns PENDING with no pending operation. In the cloud, the service rejects that output with InvalidParameterValueException: Cannot return PENDING status with no pending operations., the same error as in #736. Every replay reaches the same state, so the execution fails.
This happens in two cases.
Case 1: the branch never started.
- A completion config such as
firstSuccessful()completes the parallel before every branch has started. ParallelOperation.handleCompletionrecords each branch whose future is not complete asSKIPPED.- The SDK does not start that branch later. Nothing completes its future until the SDK suspends the invocation.
get()on the branch releases the calling thread and waits. If no other thread is active, the SDK suspends the invocation. The suspension ends the wait with the SDK's internalSuspendExecutionException.- On replay,
ParallelOperation.branch()readsSKIPPEDfrom the checkpointed result and does not queue the branch. So every replay repeats steps 3 and 4.
Case 1 needs no race, and it fails in the first invocation. The repro uses firstSuccessful(). handleCompletion applies the same rule for every completion config that can complete early: minSuccessful(n), allSuccessful(), toleratedFailureCount(n), and shouldComplete(...).
Case 2: the branch was still running when the parallel completed.
- With unlimited concurrency, branch-b can still be running when branch-a completes the parallel. The checkpointed result records branch-b as
SKIPPED. - branch-b finishes later in the same invocation.
ChildContextOperation.handleChildContextSuccessdoes not checkpoint the result, because the parallel is already complete.get()returns the result from memory. - A later operation, such as a wait, suspends the invocation. The next invocation replays the handler.
- On replay,
ParallelOperation.branch()readsSKIPPEDand does not run branch-b. Theget()call that returned a value in the first invocation now waits until the SDK suspends, as in case 1.
In case 2, the same handler code returns a value in the first invocation and fails on replay.
Steps to Reproduce
Both cases run in sdk-integration-tests with LocalDurableTestRunner.
Case 1 is the basic pattern from docs/core/parallel.md with firstSuccessful():
var runner = LocalDurableTestRunner.create(String.class, (input, context) -> {
var config = ParallelConfig.builder()
.maxConcurrency(1)
.completionConfig(CompletionConfig.firstSuccessful())
.build();
var parallel = context.parallel("first-successful", config);
var a = parallel.branch("branch-a", String.class, ctx -> "A");
var b = parallel.branch("branch-b", String.class, ctx -> "B");
parallel.get(); // MIN_SUCCESSFUL_REACHED, statuses [SUCCEEDED, SKIPPED]
a.get(); // "A"
return b.get(); // never returns a value or throws a branch exception
});
var result = runner.runUntilComplete("input");
assertEquals(ExecutionStatus.PENDING, result.getStatus());
The result has status PENDING and no error. The execution state holds first-successful (Parallel, SUCCEEDED) and branch-a (ParallelBranch, SUCCEEDED). No operation is pending. The invocation logs Invalid suspension. No operation is pending. Each further runner.run("input") gives the same result. The outcome is the same with NestingType.FLAT, and with DurableFuture.allOf(...) over both branches in place of b.get().
Case 2 uses unlimited concurrency. A wait after the parallel forces a replay:
var runner = LocalDurableTestRunner.create(String.class, (input, context) -> {
var config = ParallelConfig.builder()
.completionConfig(CompletionConfig.firstSuccessful())
.build();
var parallel = context.parallel("first-successful", config);
var a = parallel.branch("branch-a", String.class, ctx -> sleepThen(200, "A"));
var b = parallel.branch("branch-b", String.class, ctx -> sleepThen(1000, "B"));
parallel.get(); // statuses [SUCCEEDED, SKIPPED]: branch-b is still running
var value = a.get() + b.get(); // "AB" in the first invocation
context.wait("pause", Duration.ofSeconds(1)); // any later suspension causes a replay
return value;
});
runner.run("input"); // invocation 1
runner.advanceTime(); // completes "pause"
runner.run("input"); // invocation 2, a replay
sleepThen(ms, value) calls Thread.sleep(ms) and returns value.
| Invocation | a.get() + b.get() |
Output | Pending operations |
|---|---|---|---|
| 1 | returns "AB" |
PENDING |
1 (pause) |
| 2 | does not return | PENDING, no error |
0 |
| 3 and 4 | does not return | PENDING, no error |
0 |
Invocations 2 to 4 each log Invalid suspension. No operation is pending. The outcome is the same with NestingType.FLAT.
SDK Version
2.2.1. Reproduced on main at ef88276 (2.2.2-SNAPSHOT). No file under sdk/src/main changed between v2.2.1 and ef88276.
Java Version
21
Is this a regression?
Unknown
Additional Context
- #736 listed this as a follow-up: "Make
get()on a persisted skipped branch fail explicitly instead of waiting forever." #737 fixed the race in #736 and did not change this behavior. #736 described the replay of a persisted result. Case 1 also fails in the first invocation. - A fix for case 2 has to make the first invocation and the replay agree. One option:
get()throws for every branch that the checkpointed result records asSKIPPED, including a branch that finished later in the first invocation. Another option: the SDK checkpoints a late branch result so that replay returns it. - The
ParallelResulttable indocs/core/parallel.mddoes not listskipped()orstatuses(). - The end state,
PENDINGwith no pending operation, is the state that #370 and #736 also reached. How the SDK reports that state is a separate issue.
- 主要语言
- Java
- 星标
- 28
- 派生
- 13
- 平均合并
- 2 天 8 小时
- 30 天内合并 PR
- 40
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aws/aws-durable-execution-sdk-java 的其他 Issue
-
bug pkg:sdk
难度 2/5 1-3 小时 新手友好度 82/100
aws/aws-durable-execution-sdk-java#773 ·
维护者通常 1 天内回复
-
documentation pkg:sdk
难度 1/5 1-3 小时 新手友好度 88/100
aws/aws-durable-execution-sdk-java#645 · 1 条评论 ·
维护者通常 1 天内回复
-
enhancement
难度 2/5 1-3 小时 新手友好度 68/100
aws/aws-durable-execution-sdk-java#300 ·
维护者通常 1 天内回复
-
[Bug]: root handler instrumentation misses the canonical OTel execution context可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭needs-triage
难度 5/5 一周以上 新手友好度 40/100
aws/aws-durable-execution-sdk-java#770 ·
维护者通常 1 天内回复
-
[Feature]: Propagate per-operation trace context for chained invokes可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭enhancement needs-triage
难度 5/5 一周以上 新手友好度 38/100
aws/aws-durable-execution-sdk-java#764 ·
维护者通常 1 天内回复
查看 aws/aws-durable-execution-sdk-java 的全部 Issue
相似的 Issue
-
[destination-snowflake] Custom domains rejected unlike source connections可能已有人在做 @kuza55 今天认领。 未关闭autoteam community connectors/destination/snowflake team/use
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复
-
area-dashboard
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复
-
component/operate kind/feature-request
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
Forge coverage prompts carry text the agent cannot act on可能已有人在做 @graalvmbot 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 85/100
oracle/graalvm-reachability-metadata#10572 ·
维护者通常 1 天内回复
-
[CI] Core CI doesn't run for changes to amoro-format-lance (and amoro-web)可能已有人在做 @MarkAlex1234 今天认领。 未关闭
难度 1/5 1 小时以内 新手友好度 88/100
维护者通常 2 天内回复