[Feature]: Add durable terminal-scope API for cleanup and compensation
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 30/100
- issue の種類
- 機能追加
- 明瞭さ
- 説明が足りない
- 活発さ
- 活発
- 技術スタック
- java
調査の方向性
RFC/ADR の要件と提案されたライフサイクル契約から始め、次に Java SDK の現在のサスペンションモデルと、永続的な child-context または step の動作を確認します。サスペンション、リプレイ、順序付け、エラー、キャンセル、互換性に関する API とポリシーが解決され、列挙された受け入れケースがカバーされていれば完了です。
索引モデルが issue の本文から書いたものです。
説明
What would you like?
Add an SDK-supported durable terminal-scope API for registering cleanup and compensation actions that run when a scope reaches a true terminal outcome, but do not run when the current invocation merely suspends.
The exact API name and shape should be decided through an RFC/ADR. An illustrative Java API is:
return context.terminalScope(
"review-with-microvm",
ReviewResult.class,
(scopeContext, terminal) -> {
var vm = scopeContext.step(
"launch-microvm",
MicroVm.class,
stepContext -> launchMicroVm());
terminal.cleanup(
"terminate-microvm",
cleanupContext -> cleanupMicroVm(vm.id()));
terminal.compensate(
"cancel-review",
compensationContext -> cancelReview(vm.id()));
var result = scopeContext.waitForCallback(
"review-complete",
ReviewResult.class,
(callbackId, stepContext) ->
dispatchReview(vm, callbackId));
return result;
});
The important capability is the lifecycle contract, not the proposed names:
| Scope outcome | Compensation actions | Cleanup actions |
|---|---|---|
| Success | Do not run | Run |
| Application/operation failure | Run in reverse registration order | Run |
| Explicit cancellation/termination | Policy-controlled; preferably enabled by default | Run when the runtime can execute terminal work |
| Durable suspension/replay boundary | Do not run | Do not run |
Problem
Java normally encourages cleanup through try/finally or AutoCloseable. That model is unsafe around durable operations because the Java SDK currently suspends by throwing SuspendExecutionException, which unwinds the synchronous call stack. Java therefore executes active finally blocks during a normal suspension.
This pattern can release a resource while the durable execution is still logically using it:
try {
var vm = context.step(
"launch-microvm",
MicroVm.class,
stepContext -> launchMicroVm());
return context.waitForCallback(
"review-complete",
ReviewResult.class,
(callbackId, stepContext) -> dispatchReview(vm, callbackId));
} finally {
context.step(
"terminate-microvm",
Void.class,
stepContext -> terminateMicroVm());
}
When waitForCallback suspends, finally runs and terminates the MicroVM. The same issue applies to any suspending operation, including waits, wait-for-condition polling, invokes, retry delays, and suspension within map or parallel branches.
Users can manually duplicate cleanup after the success path and in catch (Exception):
try {
var result = performDurableWorkThatMaySuspend(context);
context.step("cleanup", Void.class, stepContext -> cleanup());
return result;
} catch (Exception error) {
context.step("cleanup", Void.class, stepContext -> cleanup());
throw error;
}
However, that workaround:
- duplicates orchestration code;
- becomes difficult to maintain with multiple acquired resources;
- makes reverse-order compensation cumbersome;
- is easy to implement inconsistently across nested, map, and parallel scopes;
- does not provide a clear cancellation policy;
- encourages users to reach for
finally, which has the wrong durable lifecycle semantics; - cannot express the intent as clearly as an SDK lifecycle primitive.
Goals
- Provide an idiomatic way to express work that must occur on logical completion/failure rather than invocation exit.
- Make suspension a distinct non-terminal outcome and guarantee that terminal actions do not run because of it.
- Support general orchestration scopes, not only callbacks or resource acquisition.
- Support both unconditional terminal cleanup and failure-only compensation.
- Preserve deterministic replay and stable durable-operation identity.
- Ensure terminal actions are themselves durable, retryable, observable, and replay-safe.
- Work inside top-level handlers and isolated child contexts used by map/parallel operations.
- Leave room for a language-neutral lifecycle contract that other Durable Execution SDKs can expose idiomatically.
Non-goals
- Guarantee that cleanup runs after infrastructure-level hard termination when no invocation is available to execute it.
- Replace external leases, TTLs, or reapers for resources that must eventually be reclaimed under every failure mode.
- Provide exactly-once external side effects. Cleanup and compensation steps still require normal durable-step idempotency.
- Treat suspension as failure, cancellation, or scope completion.
- Reuse
AutoCloseable, try-with-resources, or JVMfinally; those constructs are tied to stack unwinding, not durable terminal state.
Possible Implementation
1. Scope and registration model
terminalScope could execute a deterministic body against a child durable context and provide a registration object:
<T> T terminalScope(
String name,
Class<T> resultType,
DurableTerminalScopeFunction<T> body);
Illustrative registration API:
interface DurableTerminalActions {
void cleanup(String name, DurableTerminalAction action);
void compensate(String name, DurableTerminalAction action);
}
Possible future overloads could accept configuration for action ordering, retry policies, cancellation behavior, and cleanup-error aggregation.
Registrations should be deterministic declarations. On each replay, execution reruns the scope body and reconstructs the same registrations before reaching the same suspension or terminal path. Registration names and ordering must remain stable for a given checkpoint history.
2. Suspension handling
With the current synchronous Java execution model, the scope implementation can distinguish internal suspension from a terminal failure:
try {
T result = body.run(childContext, actions);
runCleanup(actions);
return result;
} catch (SuspendExecutionException suspension) {
// Suspension is not terminal. Run no compensation or cleanup.
throw suspension;
} catch (Exception failure) {
runCompensation(actions);
runCleanup(actions);
throw failure;
}
This is only conceptual pseudocode. The implementation must preserve all SDK control-flow errors and should not accidentally convert or swallow Error instances. The final design should use an internal outcome classifier rather than encourage broad user-visible catch (Throwable).
3. Durable identity and replay
The scope should own stable namespaces for:
- body operations;
- each registered compensation;
- each registered cleanup;
- terminal phase/progress if explicit checkpointing is required.
Terminal actions must execute through durable child contexts or equivalent SDK-managed operations. Completed actions must consume checkpoints on replay and must not repeat their bodies.
The design must define behavior when:
- the body suspends before all registrations are reached;
- the body fails after acquiring several resources;
- a compensation action suspends;
- a cleanup action suspends;
- a compensation or cleanup action fails and is retried;
- replay observes some terminal actions completed and later actions pending;
- a scope is nested inside another terminal scope;
- a scope is used within each item of a map or branch of a parallel operation.
Registration should not rely only on ephemeral in-memory state after a terminal phase begins. Either deterministic replay must reconstruct the complete registration set before resuming terminal work, or the SDK must durably record sufficient scope metadata to resume it safely.
4. Ordering
Suggested defaults:
- compensations run in reverse registration order, matching the usual acquisition/rollback model;
- cleanups run in reverse registration order so resources unwind like a stack;
- cleanup still runs if a compensation fails, where execution policy permits;
- multiple failures are preserved through a documented primary/suppressed or aggregate-error model.
Configuration could later permit forward or parallel execution, but the initial API should favor deterministic sequential behavior.
5. Error semantics
The RFC should specify:
- whether all terminal actions are attempted after one fails;
- how the original body failure is preserved;
- how cleanup/compensation failures are attached or aggregated;
- which retry configuration applies to each action;
- whether a failed cleanup changes an otherwise successful scope into a failed scope;
- what happens if terminal actions exhaust retries;
- how cancellation arriving during terminal processing is handled.
A reasonable initial policy is:
- on body failure, preserve the body failure as primary;
- attempt remaining compensation and cleanup actions;
- attach terminal-action failures as suppressed/aggregate details;
- on body success, a cleanup failure fails the scope;
- allow per-action retry configuration using existing step retry semantics.
6. Cancellation and hard termination
Cancellation must be defined separately from suspension.
The API could expose a policy such as:
TerminalScopeConfig.builder()
.compensateOnCancellation(true)
.cleanupOnCancellation(true)
.build();
The contract must acknowledge that terminal actions can run only when the Durable Execution service schedules code to process the terminal transition. For externally terminated compute or lost invocations, users still need resource-native leases, TTLs, idempotent deletion, or a reaper as a safety net.
7. Observability
Execution history and telemetry should make the lifecycle explicit:
- scope started;
- scope suspended without terminal actions;
- terminal outcome selected;
- compensation started/completed/failed;
- cleanup started/completed/failed;
- scope terminal processing completed.
This is important for diagnosing cleanup that is delayed, retried, or partially completed.
8. Compatibility and rollout
This can be introduced as an additive API, so it should not be a breaking change. Because it creates a new durable orchestration primitive and potentially a cross-SDK lifecycle contract, it should go through an RFC/ADR and conformance review.
If a new operation/checkpoint type is introduced, backward compatibility and older-runtime behavior must be specified. If implemented initially as SDK composition over existing child contexts and steps, operation identity and upgrade behavior still require tests.
9. Testing and acceptance criteria
- Suspension from each suspending operation type runs no registered terminal action.
- Replay reconstructs registrations deterministically.
- Success runs cleanup exactly through durable/checkpointed semantics.
- Failure runs compensation and cleanup in the documented order.
- Cancellation follows an explicit, tested policy.
- A terminal action can itself suspend and resume.
- Completed terminal actions are skipped on replay.
- Terminal-action retries do not repeat already completed actions.
- Nested terminal scopes behave deterministically.
- Map/parallel branches use isolated scope contexts and do not interfere.
- Multiple body/compensation/cleanup failures preserve documented error information.
- API docs warn that cleanup side effects remain at-least-once and should be idempotent.
- Examples include the MicroVM lifecycle use case.
- Java error-handling documentation cross-links the feature when available.
- Language-neutral requirements are proposed if this becomes a cross-SDK capability.
10. Open design questions
- Should
cleanuprun on success and failure whilecompensateruns only on failure/cancellation, or should outcome-specific hooks be exposed directly? - Is registration-before-acquisition required, or may a resource be acquired and then registered?
- Must registrations be checkpointed when declared, or is deterministic reconstruction sufficient?
- Should terminal actions receive the body result or failure?
- Should actions be Java lambdas captured during replay, serializable descriptors, or named operations reconstructed by user code?
- Should terminal processing be a new service-visible operation type or SDK composition over existing primitives?
- What is the cancellation contract supported by the service today?
- How should cleanup behave when a map uses early completion and abandons unfinished branches?
- Should compensation be included in the first version or added after a cleanup-only terminal scope?
- What naming best distinguishes durable terminal lifecycle from JVM lexical scope?
Is this a breaking change?
No. The proposal is an additive API.
Does this require an RFC?
Yes. It introduces new lifecycle semantics, replay rules, error aggregation, cancellation policy, and likely cross-SDK considerations.
Additional Context
The proposed terminalScope shape is a synthesis for Durable Execution rather than a direct copy of another SDK. Related established patterns include:
- the Saga pattern and reverse-order compensation;
- Temporal Java's
Saga.addCompensation(...)/compensate(); - Cadence Java's similar Saga helper;
- workflow-engine compensation handlers such as BPMN compensation;
- asynchronous resource APIs such as Reactor
usingWhen, which distinguish completion, error, and cancellation cleanup but are not durable/replay-aware.
The durable-specific requirement is to distinguish suspension from all terminal outcomes. A normal invocation boundary must not trigger cleanup or compensation.
Related issues:
- Java documentation for the current
finallyhazard: https://github.com/aws/aws-durable-execution-sdk-java/issues/645 - Cross-language documentation pattern: https://github.com/aws/aws-durable-execution-docs/issues/131
- Detailed cross-language discussion: https://github.com/aws/aws-durable-execution-docs/issues/131#issuecomment-5399335840
- 主要言語
- Java
- スター
- 28
- フォーク
- 13
- 平均マージ
- 2日 8時間
- マージ済み PR(30日)
- 40
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
aws/aws-durable-execution-sdk-java のほかの issue
-
bug pkg:sdk
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
aws/aws-durable-execution-sdk-java#773 ·
メンテナーはふだん 1 日以内に返信
-
documentation pkg:sdk
難易度 1/5 1〜3時間 初心者へのやさしさ 88/100
aws/aws-durable-execution-sdk-java#645 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
aws/aws-durable-execution-sdk-java#300 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: root handler instrumentation misses the canonical OTel execution context対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープンneeds-triage
難易度 5/5 1週間以上 初心者へのやさしさ 40/100
aws/aws-durable-execution-sdk-java#770 ·
メンテナーはふだん 1 日以内に返信
-
[Feature]: Propagate per-operation trace context for chained invokes対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープンenhancement needs-triage
難易度 5/5 1週間以上 初心者へのやさしさ 38/100
aws/aws-durable-execution-sdk-java#764 ·
メンテナーはふだん 1 日以内に返信
aws/aws-durable-execution-sdk-java の issue をすべて見る
似ている issue
-
[destination-snowflake] Custom domains rejected unlike source connections対応中かも @kuza55 が今日担当しました。 オープンautoteam community connectors/destination/snowflake team/use
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
area-dashboard
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
component/operate kind/feature-request
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
Forge coverage prompts carry text the agent cannot act on対応中かも @graalvmbot が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
oracle/graalvm-reachability-metadata#10572 ·
メンテナーはふだん 1 日以内に返信
-
[CI] Core CI doesn't run for changes to amoro-format-lance (and amoro-web)対応中かも @MarkAlex1234 が今日担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 2 日以内に返信