[Feature]: Support CompletionStage-returning durable step functions
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- java
- Domain
- api, backend, distributed-systems
Research direction
Start by reading the existing stepAsync overloads, DurableFuture handling, and execution manager's active-thread tracking. Use LocalDurableTestRunner to exercise delayed completion, replay, failures, timeout, retry, cancellation, null stages, and suspension races. Done means CompletionStage-backed steps preserve checkpointing, context and lifecycle behavior while existing synchronous APIs remain compatible.
Written by the indexing model from the issue text.
Description
What would you like?
Add a supported API for durable step functions whose user function returns a CompletionStage<T> (typically a CompletableFuture<T>), so applications can use asynchronous Java clients without blocking an SDK user-executor thread.
An illustrative API is:
DurableFuture<Response> response = context.stepStage(
"fetch",
Response.class,
stepContext -> asyncClient.sendAsync(request));
CompletionStage should be the public abstraction rather than requiring the concrete CompletableFuture type. The exact name is open to design; a distinct method such as stepStage avoids erasure and overload ambiguity with the existing Function<StepContext, T> methods.
This is intended as a targeted addition, not a proposal to make the entire SDK CompletableFuture-first. The existing synchronous APIs, DurableFuture, and virtual-thread executor support should remain.
Benefits
- Direct async-client integration: Applications can return stages from AWS SDK async clients, async HTTP clients, database drivers, and other non-blocking libraries without calling
.join()inside a step. - Better Java 17 scalability: The SDK targets Java 17, where virtual threads are unavailable. A stage-returning step avoids holding one platform thread for each in-flight asynchronous request.
- Lower per-operation overhead at high concurrency: Virtual threads are much cheaper than platform threads, but each pending virtual thread still carries stack and thread-local state. Tracking a stage directly can reduce that overhead for very large fan-out workloads.
- Clear ownership of lifecycle semantics: The SDK can track the stage as active in-process work, checkpoint its terminal result, normalize failures, and preserve retry/replay behavior instead of requiring every application to build a blocking adapter.
- Complementary to virtual threads: Virtual threads remain the simplest option for blocking libraries, while stage-returning steps provide a natural path for libraries that are already asynchronous.
This feature does not reduce Lambda compute time for an in-flight network request: the invocation must remain active until the stage completes because the request and continuation are in-memory. Durable waits already handle compute-free suspension. The benefit here is avoiding a blocked user thread while preserving the same invocation-lifecycle behavior.
Example use cases
- Call
DynamoDbAsyncClient,S3AsyncClient,LambdaAsyncClient, or another AWS async client from a durable step. - Use Java 11+
HttpClient.sendAsync()or a third-party async HTTP client for remote API calls. - Compose several asynchronous calls inside one checkpointed step with
thenCompose/thenCombine, then checkpoint the final value. - Run high-fan-out I/O workloads on Java 17 without configuring a large platform-thread pool.
- Integrate third-party libraries that expose
CompletionStageas their primary API without converting them to blocking calls.
Possible Implementation
Add stepStage overloads parallel to the existing stepAsync overloads:
<T> DurableFuture<T> stepStage(
String name,
Class<T> resultType,
Function<StepContext, ? extends CompletionStage<T>> function);
<T> DurableFuture<T> stepStage(
String name,
TypeToken<T> resultType,
Function<StepContext, ? extends CompletionStage<T>> function,
StepConfig config);
The RFC should determine the final naming and whether a dedicated functional interface would produce a clearer API.
Suggested lifecycle:
- Start the step using the same operation identity, replay lookup, retry policy, plugin hooks, and serialization rules as a synchronous step.
- Invoke the user function with an attached
StepContext. - Validate that the returned stage is non-null.
- Treat the pending stage as active in-process work even though no user thread is blocked. The execution must not suspend while the stage is pending, because its request and continuation would be lost with the invocation.
- On completion, run SDK-owned terminal processing with the required context attached, normalize the value or exception, and checkpoint it through the existing step machinery.
- On replay of a completed step, return the checkpointed result without invoking the user function or creating a new stage.
Design considerations:
- The current execution manager tracks active threads. Supporting this safely may require an active-work token or equivalent mechanism for asynchronous work that is not represented by a live SDK user thread.
- Completion can occur on an arbitrary client/event-loop thread. SDK checkpointing, plugin hooks, logging metadata, OpenTelemetry context, and thread-local SDK context must be propagated or restored deliberately.
- Synchronously thrown exceptions and exceptional stage completion should have identical retry and step-failure semantics.
- Step timeout behavior should cover the lifetime of the returned stage.
- Cancellation semantics must be documented. Cancelling a local stage may be best-effort and is not equivalent to cancelling a checkpointed durable operation.
- User-defined completion callbacks are ordinary in-process step code and are not individually checkpointed. Only the final step result is durable.
- The implementation should not use the common fork-join pool implicitly for SDK terminal processing.
Acceptance criteria:
- A durable step can return a
CompletionStage<T>and still expose its operation result asDurableFuture<T>. - Already-completed and asynchronously-completed stages both checkpoint successfully.
- Exceptional and cancelled stages produce documented failure/retry behavior.
- The execution cannot suspend early while a stage-backed step has pending in-process work.
- Completed steps replay without invoking the asynchronous function again.
- Existing step retry, timeout, serialization, and per-retry/per-run semantics apply consistently.
- SDK context, logging metadata, and plugin/OpenTelemetry lifecycle behavior are correct when completion originates on a non-SDK thread.
- Unit tests cover success, failure, cancellation, timeout, retry, replay, null stages, and suspension races.
- An integration test uses
LocalDurableTestRunnerto demonstrate delayed asynchronous completion and replay. - Documentation includes an async HTTP or AWS SDK async-client example and explains when to choose stages versus virtual threads.
- Existing synchronous and
DurableFutureAPIs remain source- and binary-compatible.
Is this a breaking change?
No
Does this require an RFC?
Yes
Additional Context
The SDK already uses CompletableFuture internally for operation coordination, and DurableFuture.get() integrates blocking waits with active-thread deregistration and durable suspension. This proposal is specifically about accepting asynchronous user work as a step body while retaining SDK-controlled checkpoint/replay semantics.
A raw CompletableFuture<T> for a durable operation should not be exposed as though it survives invocation suspension. The returned public operation handle should remain DurableFuture<T>; the CompletionStage<T> is only the in-process implementation of the step body.
Virtual threads and this feature solve related but different problems:
- Virtual threads make blocking code scalable on Java 21+.
- Stage-returning steps integrate naturally with already-asynchronous libraries and also improve scalability on Java 17.
- Durable suspension remains responsible for stopping an invocation when no in-process work can make progress.
- Dominant language
- Java
- Stars
- 28
- Forks
- 13
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 40
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from aws/aws-durable-execution-sdk-java
-
bug pkg:sdk
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
aws/aws-durable-execution-sdk-java#773 ·
Maintainers usually reply within 1 day
-
documentation pkg:sdk
Difficulty 1/5 1-3 hours Newbie friendliness 88/100
aws/aws-durable-execution-sdk-java#645 · 1 comment ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
aws/aws-durable-execution-sdk-java#300 ·
Maintainers usually reply within 1 day
-
[Bug]: root handler instrumentation misses the canonical OTel execution contextPossibly taken A pull request linked to this issue is open or already merged. Openneeds-triage
Difficulty 5/5 Over a week Newbie friendliness 40/100
aws/aws-durable-execution-sdk-java#770 ·
Maintainers usually reply within 1 day
-
[Feature]: Propagate per-operation trace context for chained invokesPossibly taken A pull request linked to this issue is open or already merged. Openenhancement needs-triage
Difficulty 5/5 Over a week Newbie friendliness 38/100
aws/aws-durable-execution-sdk-java#764 ·
Maintainers usually reply within 1 day
All issues in aws/aws-durable-execution-sdk-java
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
apache/skywalking#14120 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
HMCL-dev/HMCL#6943 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
micronaut-projects/micronaut-core#13677 ·
Maintainers usually reply within 1 day
-
new feature
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/rocketmq-dashboard#5594 ·
Maintainers usually reply within 3 days