[Feature]: Add LMI cloud end-to-end tests for concurrency and invocation lifecycle
メンテナーはふだん 1 日以内に返信
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 25/100
調査の方向性
The existing E2E workflow is in .github/workflows/e2e-tests.yml. The CloudDurableTestRunner in sdk-testing/src/main/java/software/amazon/lambda/durable/testing/ is the driver. Start by understanding the current test deployment and the LMI-specific infrastructure requirements. 'Done' means a new LMI test suite is integrated into CI, with evidence of concurrent invocations in the same JVM and coverage of the lifecycle scenarios from #726.
索引モデルが issue の本文から書いたものです。
説明
What would you like?
Add a dedicated end-to-end cloud test suite for the Java Durable Execution SDK on Lambda Managed Instances (LMI), including PerExecutionEnvironmentMaxConcurrency > 1.
This is the cloud-validation follow-up to #726. That issue contains local reproductions of unbounded invocation cleanup, a PENDING response preceding root-handler cleanup, and starvation when concurrent handlers share a fixed executor. Real LMI tests are needed to validate the runtime behavior, the eventual fixes, and future regressions.
The existing E2E workflow deploys examples and runs CloudBasedIntegrationTest with a Java 17/21/25 matrix and JUnit parallelism. It does not configure an LMI-specific deployment or prove that concurrent invocations share one Java execution environment. Driver-side parallelism alone cannot exercise that requirement.
LMI-specific behavior to cover:
- Multiple OS runtime threads share the handler object and process-wide resources.
- Execution environments remain active between invocations.
- An invocation timeout does not forcibly terminate the function's code, so task cancellation, cleanup, and worker-slot recovery require explicit validation.
References: Java runtime for LMI and LMI execution environment lifecycle.
Possible Implementation
1. Deploy a dedicated LMI test stack.
Add repeatable infrastructure and small purpose-built handlers, following the existing example/cloud-test structure where practical:
- A test capacity provider and required networking/roles, or explicit parameters for an existing dedicated test capacity provider.
- Durable-enabled function versions/aliases configured for LMI, with the deployed version, runtime, capacity-provider association, and
PerExecutionEnvironmentMaxConcurrencyrecorded in test artifacts. - Separate fixtures/configurations for concurrency 1, 2, and a modest higher value such as 8. Treat those as LMI invocation concurrency values, independently of SDK map/parallel branch concurrency.
- Short, configurable invocation timeouts for deadline tests, distinct from the overall durable execution timeout.
- Explicit test-run identifiers and isolated resources for coordination, side-effect observation, and logs.
Verify LMI + Durable Execution availability and supported Java runtime/architecture combinations in the chosen test region before running the matrix. Use only supported combinations; the current ordinary-Lambda Java matrix should not be copied blindly. An unavailable deployment combination must be reported explicitly, without falling back to ordinary Lambda and claiming LMI coverage.
Keep the test SDK artifact tied to the commit under test.
2. Build a driver that proves concurrent execution in the same environment.
Reuse CloudDurableTestRunner and its history parsing where appropriate. Create one runner per execution because it stores mutable lastResult; share the thread-safe AWS Lambda client if useful.
Add test-only diagnostics that record:
- A JVM/execution-environment identifier generated once during initialization.
- Durable execution ARN, invocation request ID, and test-run ID.
- Runtime-handler, root-handler, and step/child task entry/exit markers.
- Executor active/queued task counts and invocation-owned task counts where available.
- Root cleanup entry/exit and the point where the outer wrapper returns the SDK response.
- Same-JVM sequence numbers or monotonic timestamps for ordering assertions.
Keep diagnostic identifiers out of durable operation names, control flow, and checkpoint identity.
Use controlled, bounded barriers to establish overlapping invocations, then require evidence of distinct invocation IDs overlapping within the same environment ID. Setting concurrency to 2, sending two requests, or seeing overlapping CloudWatch timestamps from different environments is insufficient.
Placement is service-controlled. Use a bounded warm-up/admission procedure with a clear placement deadline; if same-environment overlap cannot be established, fail/report an infrastructure-precondition failure rather than passing the concurrency case. Do not assume a small capacity provider guarantees one JVM.
For worker-recovery assertions, keep the other established slots occupied with bounded healthy work, then require a follow-up probe to enter the affected environment. Report environment replacement separately: a probe succeeding in a new environment does not prove that the old invocation released its worker. For higher-concurrency configurations, verify restored concurrent admission in the original environment rather than only one successful follow-up request.
Collect ordering at the actual execution boundaries. In particular, the current onInvocationEnd hook is not proof that the root handler has exited or that the runtime response has been returned. A test wrapper around the SDK entry point plus root/task finally markers can establish that order.
3. Cover the following scenarios against the real service.
| Scenario | Required assertions |
|---|---|
| Concurrency-1 baseline | A step/wait/resume execution completes correctly on LMI and establishes baseline diagnostics. |
| Overlapping executions with identical operation sequences | Two or more invocations in one JVM use different input/result markers. Results, failures, checkpoint histories, and invocation-owned tasks remain attributable to the correct execution. |
| Suspension and root cleanup (#726) | A root handler suspends inside a bounded try/finally fixture. Root cleanup exits before the outer wrapper returns a normal PENDING response. Another invocation in that same JVM continues to make progress. Use invocation-local cleanup instrumentation, without adding durable operations to finally. |
| Resume/replay under concurrent load | A completed step is followed by a durable wait or callback, then resumes while other executions are active. Invocation IDs demonstrate an actual resume; completed work is replayed without re-running its body, and stored results/failures and operation identity remain stable. Do not require the resumed invocation to land in the original JVM. |
| Real invocation timeout and cooperative cancellation (#726) | Use an invocation-owned, deliberately blocked but interruptible operation. Exercise the real LMI invocation timeout/deadline path. Observe the invocation outcome, actual task exit, cleanup budget, and availability of the affected worker capacity afterward. A co-located healthy invocation continues and subsequent probes demonstrate recovery. |
| Timeout/cleanup limitations | A separately bounded fixture that ignores interruption documents the unsupported case without hanging the suite indefinitely. Do not assert that Java can forcibly stop arbitrary user code; record residual activity and require explicit diagnostics and prevention of further SDK work under the eventual #726 contract. |
| Shared fixed executor and nested orchestration (#726) | Two synchronized invocations share a two-thread user executor and each awaits a step; also exercise nested child contexts/map/parallel. Supported configurations must progress, or explicitly unsupported configurations must be rejected with the documented diagnostic. The current implementation should reproduce the regression rather than appear successful through infrastructure replacement. |
| In-flight work and failure cleanup | A handler returns or fails while a child/step is still active. Required tasks drain/cancel within budget, pending SDK cleanup is attempted, and the healthy invocation's executor/tasks remain usable. Assert no late SDK checkpoint activity after the closing boundary defined by the fix. |
| Repeated warm-environment execution | Run bounded batches of success, failure, suspend/resume, and timeout cases. After a defined quiescence interval, invocation-owned tasks return to baseline and worker capacity remains available; queued work and retained completed-invocation context do not accumulate. Allow documented cached-pool idle retention rather than requiring zero JVM threads immediately. |
Use real Durable Execution checkpoint/history APIs and real waiting/resumption. Local fake backends and time-skipping runners remain useful unit/integration coverage, but do not satisfy these cloud cases.
Distinguish invocation-level timeout/failure from the overall durable execution outcome. A durable execution may be retried/resumed by the service; final success alone must not hide a timed-out invocation whose old tasks are still running. Distinguish client HTTP timeouts from server-reported invocation timeouts as well.
For replay assertions, use durable history and a test-only attempt ledger with idempotent business effects. Require an already-checkpointed successful/failed step body to be skipped on replay. For work interrupted before its checkpoint, allow the retries permitted by the configured step semantics; do not assert exactly-once execution where the SDK guarantees at-least-once behavior.
4. Integrate with CI and preserve useful failure evidence.
Add a separate LMI workflow/job, with explicit opt-in for local cloud tests and PR runs. Run the bounded suite through manual dispatch and automatically on every push to main, including every merged change, using the existing test-account/OIDC model. Do not filter main-branch runs by changed paths.
- Serialize updates to shared stacks/capacity providers, or use distinct resources per run.
- Give provisioning, waiting for prior executions, placement, test execution, and history/log collection separate finite time budgets.
- Bound request counts, capacity, retries, and total test duration; ensure blocking fixtures have a test-only escape path.
- Always publish JUnit reports, deployed configuration/commit, environment/request/execution correlation, durable histories, thread/task snapshots, and relevant CloudWatch logs.
- Use event ordering and bounded polling to handle eventual log/history delivery; avoid tests that pass or fail solely on a fixed sleep.
- Report deployment/placement failures separately from assertion failures. Once a scenario has been established, a violated lifecycle assertion must not be retried until it turns green.
- Reuse one persistent test stack and its LMI functions across runs. Keep the stack and functions after test failures or timeouts; do not add automatic resource teardown or a janitor. Document ownership of the persistent stack, bucket, and capacity provider. Isolate and expire per-run control data, bound log retention, and serialize code updates with tests so unfinished executions are not resumed on a different SDK build.
Acceptance Criteria
- Infrastructure can reproducibly deploy a real LMI durable function using the SDK commit under test.
- The recorded matrix includes concurrency 1 and at least one value greater than 1, using supported runtime/region combinations.
- Every same-environment concurrency case includes positive evidence of overlapping invocation IDs in one JVM.
- Cloud regressions cover all three core lifecycle/executor problems in #726.
- Tests assert the desired fixed behavior: they expose the affected implementation and pass after the relevant fixes, rather than counting reproduced defects as successful regression coverage.
- Real invocation timeout, task exit, cleanup budget, healthy-invocation isolation, and worker-capacity recovery are measured separately.
- At least one real suspension/resume path verifies deterministic replay, stored successes/failures, and appropriate side-effect counts.
- Default executor behavior is covered; custom fixed/virtual-thread variants are included only with explicit runtime support and executor contracts.
- Repeated warm-environment runs detect residual invocation tasks/context and resource growth with bounded tolerances.
- CI exposes deployment/precondition failures, assertion failures, skipped/unsupported combinations, and evidence-collection failures distinctly.
- Every push to
maintriggers the LMI cloud test suite, regardless of which paths changed. - Failure artifacts are available for manual and main-branch runs, and the same persistent stack/functions are reused without automatic resource deletion or a janitor.
- Documentation explains prerequisites, deployment/update, local driver commands, budgets, interpreting results, and ownership/retirement of persistent resources.
Is this a breaking change?
No. This adds test infrastructure and regression coverage. Any SDK behavior or executor contract changes remain tracked in #726.
Does this require an RFC?
No public API RFC is required for the initial test suite. Document a short test design covering same-environment placement, lifecycle observation, timeout semantics, and infrastructure ownership before implementation.
Additional Context
Related: #726.
The existing local investigation established the three defects on v2.2.0. No LMI cloud validation has been performed as part of that report, so this issue defines the work needed to close that gap.
This suite belongs in the Java SDK repository as runtime-specific cloud integration coverage. If implementation exposes new language-neutral requirements, propose those separately in the shared conformance repository.
- 主要言語
- Java
- スター
- 28
- フォーク
- 13
- 平均マージ
- 2日 3時間
- マージ済み PR(30日)
- 44
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
aws/aws-durable-execution-sdk-java のほかの issue
-
documentation pkg:sdk
難易度 1/5 1〜3時間 初心者へのやさしさ 88/100
aws/aws-durable-execution-sdk-java#645 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
aws/aws-durable-execution-sdk-java#300 ·
メンテナーはふだん 1 日以内に返信
-
enhancement needs-triage
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
aws/aws-durable-execution-sdk-java#779 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: root handler instrumentation misses the canonical OTel execution context対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープンneeds-triage
難易度 5/5 1週間以上 初心者へのやさしさ 40/100
aws/aws-durable-execution-sdk-java#770 ·
メンテナーはふだん 1 日以内に返信
-
[Feature]: Propagate per-operation trace context for chained invokes対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープンenhancement needs-triage
難易度 5/5 1週間以上 初心者へのやさしさ 38/100
aws/aws-durable-execution-sdk-java#764 ·
メンテナーはふだん 1 日以内に返信
aws/aws-durable-execution-sdk-java の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
-
Make branch and label autocomplete matching locale-independent対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 83/100
jenkinsci/gitlab-plugin#1950 ·
-
It's not necessary to copy the memory block in the readWrite() of org.h2.store.fs.mem.FileMemDataオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
h2database/h2database#4435 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
micronaut-projects/micronaut-core#13717 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
ADORSYS-GIS/token-status-link#145 ·
メンテナーはふだん 3 日以内に返信