[Feature]: Adaptive checkpoint batching with separate blocking and nonblocking delays
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
調査の方向性
Start with DurableConfig.java, BaseDurableOperation.java, CheckpointManager.java, and ApiRequestDelayedBatcher.java, which the issue links as the current implementation. Trace how acknowledgement requirements reach checkpoint submission and how the batcher selects and advances deadlines; then review timer tests and polling behavior. Done means deterministic tests cover the stated timing, ordering, limits, failure, shutdown, replay, and compatibility criteria, with documentation and repeatable sequential and concurrent benchmarks.
索引モデルが issue の本文から書いたものです。
説明
What would you like?
Make checkpoint batching adapt to whether an update blocks SDK progress. Reuse the existing delay setting for blocking updates and add only one new setting for nonblocking updates:
- Blocking delay: use the existing
DurableConfig.withCheckpointDelay(Duration)setting when the batch contains an update whose acknowledgement is required before the operation can continue. Its default remains zero. - Nonblocking delay: add a separate setting for a longer collection window when all updates in the batch can be persisted asynchronously.
This would let the SDK combine a fast step's asynchronous START and synchronous SUCCEED without making the caller wait for the full nonblocking window after the step finishes.
Current behavior
The SDK has one DurableConfig.withCheckpointDelay(Duration) setting, defaulting to zero. CheckpointManager applies it to ordinary operation updates regardless of whether the caller waits. The synchronous send method calls the asynchronous send method and then joins its future; the batcher is not told that this update blocks progress.
With a zero delay, START and SUCCEED can become separate API requests. Setting a longer delay improves their opportunity to batch, but also delays synchronous updates. For example, with a 100 ms delay, a START submitted at t=0 and SUCCEED submitted at t=10 ms normally wait until approximately t=100 ms to be sent together.
Requested behavior
-
Keep
withCheckpointDelay(Duration)as the blocking delay, defaulting to zero. Add only a nonblocking delay setting, so the two delays can be configured independently.withNonblockingCheckpointDelaybelow is a proposed new method;withCheckpointDelayalready exists:DurableConfig.builder() .withCheckpointDelay(Duration.ZERO) // existing setting; still the default .withNonblockingCheckpointDelay(Duration.ofMillis(100)) .build(); -
A batch containing only nonblocking updates may use the longer delay. Once a blocking update joins, bring its flush deadline forward to the earlier of the existing deadline and
blockingUpdateArrival + blockingDelay. -
Later updates must never push an existing deadline back. The default zero blocking delay requests prompt flushing as soon as a blocking update joins.
-
Preserve batching of already available updates, their order, and request size/count limits. Shortening the deadline must not force every update into its own API request.
-
Define blocking by the checkpoint's acknowledgement requirement. An at-least-once-per-retry START can be nonblocking; an at-most-once-per-retry START and a step's synchronous completion are blocking. The public use of
stepAsync()does not by itself make all of its internal checkpoints nonblocking.
With the example settings above, START at t=0 opens a window ending at t=100 ms. SUCCEED at t=10 ms requests flushing immediately, so both updates can be sent together without waiting for the rest of the 100 ms window. A caller can still set withCheckpointDelay(Duration.ofMillis(1)) to allow a short blocking collection delay instead. These delays govern intentional collection time; they do not bound thread scheduling, waiting behind an in-flight batch, or service latency.
Possible Implementation
Carry the update's acknowledgement requirement through the operation and execution layers into CheckpointManager, and select the appropriate delay when submitting to ApiRequestDelayedBatcher. The batcher already supports choosing the earliest deadline; ensure a newly urgent update also advances an already scheduled timer.
Compatibility proposal:
- Keep
withCheckpointDelay(Duration)andgetCheckpointDelay()as the blocking-delay API, with the existing default of zero. Do not introduce a separate blocking-delay setter or getter. - Add only the nonblocking delay setting and its accessor. For compatibility, when it is unset, use the configured blocking delay as its effective value; existing configurations therefore retain their current timing until the new setting is used.
- An explicit nonblocking delay affects only nonblocking updates and never overrides
withCheckpointDelay, independent of builder call order. The 100 ms nonblocking delay above is an opt-in example, not a proposed silent default change. - Validate nonnegative durations and
blockingDelay <= nonblockingDelay, and document the unset nonblocking setting's fallback. - Preserve checkpoint-token ordering, service acknowledgement before releasing blocking callers, replayed results and failures, and shutdown/failure settlement of pending futures. Keep polling schedules separate: the longer nonblocking window must not postpone a due refresh.
Acceptance criteria
- Deterministic timer tests cover nonblocking-only batches, blocking-only batches, and a blocking update joining an existing long-window batch. Use a controllable clock/scheduler rather than narrow wall-clock assertions.
- Verify that later nonblocking updates cannot postpone a blocking batch, earlier deadlines remain earlier, and zero/equal/custom delay combinations behave as documented.
- Verify that
withCheckpointDelayremains the blocking setting with a zero default, the new nonblocking setting is independent of builder call order, and leaving it unset preserves legacy configuration behavior. - Verify START/SUCCEED coalescing and blocking at-most-once-per-retry START acknowledgement before user code runs.
- Cover mixed concurrent producers, batch limits, polling, checkpoint failures, shutdown, and replay without repeating completed user work.
- Add a repeatable benchmark for 1,000 short sequential steps, alongside a concurrent-step case. Record total execution time and checkpoint request counts over repeated runs with the same Lambda configuration. Existing concurrent-only performance tests do not reliably expose per-step delay accumulated by a sequential workflow.
- Document the two settings, compatibility behavior, and latency/request-count trade-off.
Is this a breaking change?
No. Reuse the existing blocking-delay API and add only the nonblocking setting, preserving existing defaults and legacy configurations as described above.
Does this require an RFC?
Yes, to agree on the new scheduling contract, configuration names, and compatibility behavior before implementation.
Additional Context
Current implementation inspected at ac6a6f5ea50d6a9098b716905f01c609e4c201ae:
- Single checkpoint delay and default
- Synchronous send waits on the asynchronous future
- Checkpoint submission
- Batch deadline selection
Related: Python PR #725 shortens empty-queue waits when a synchronous checkpoint is present and includes sequential-step benchmarks across Python, Java, and JavaScript. It motivates adapting to blocked callers; Java can retain its own scheduling model, reuse withCheckpointDelay for blocking updates, and add a separately configurable nonblocking delay.
- 主要言語
- Java
- スター
- 28
- フォーク
- 13
- 平均マージ
- 2日 3時間
- マージ済み PR(30日)
- 44
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
aws/aws-durable-execution-sdk-java のほかの issue
-
documentation pkg:sdk
難易度 1/5 1〜3時間 初心者へのやさしさ 88/100
aws/aws-durable-execution-sdk-java#645 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
aws/aws-durable-execution-sdk-java#300 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: root handler instrumentation misses the canonical OTel execution context対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープンneeds-triage
難易度 5/5 1週間以上 初心者へのやさしさ 40/100
aws/aws-durable-execution-sdk-java#770 ·
メンテナーはふだん 1 日以内に返信
-
[Feature]: Propagate per-operation trace context for chained invokes対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープンenhancement needs-triage
難易度 5/5 1週間以上 初心者へのやさしさ 38/100
aws/aws-durable-execution-sdk-java#764 ·
メンテナーはふだん 1 日以内に返信
-
bug needs-triage
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
aws/aws-durable-execution-sdk-java#763 ·
メンテナーはふだん 1 日以内に返信
aws/aws-durable-execution-sdk-java の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
commonmark/commonmark-java#460 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
GoogleCloudPlatform/spring-cloud-gcp#4664 ·
メンテナーはふだん 1 日以内に返信
-
enhancement user story
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100