Local activity scheduleToClose budget restarts when a timer-backed retry runs on replay
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 68/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- java
- Lĩnh vực
- backend, distributed-systems
Hướng nghiên cứu
Start in temporal-sdk/src/main/java/io/temporal/internal/statemachines/LocalActivityCallback.java, tracing how firstSkd is parsed and passed into retries. Reproduce with sticky queue scheduling disabled and the listed local activity options; done means replayed retries preserve the original schedule-to-close budget, terminate with RETRY_STATE_TIMEOUT, and a regression test covers the behavior.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Expected Behavior
A local activity that retries through a workflow timer keeps counting ScheduleToCloseTimeout from its first attempt across replay. sdk-core restores the original schedule time from the marker and hands it back with the retry (local_activity_state_machine.rs:147, :646); its proto says it "Must be passed with attempt to the retry LA" (activity_result.proto:97-99). sdk-go sets the schedule time from workflow time (workflow.go:1281) and computes the deadline from it (internal_event_handlers.go:892-894), so both preserve the original scheduling baseline across replay.
Actual Behavior
Java parses firstSkd from the marker into LocalActivityFailedException, then discards it (LocalActivityCallback.java:29-37). The retry uses System.currentTimeMillis() captured when workflow code ran, which on replay is the replay clock. The retry pre-check then sees a budget computed from that clock, so a replay can reset the budget. In the reproduction below the activity runs all 5 attempts and ends with RETRY_STATE_MAXIMUM_ATTEMPTS_REACHED instead of RETRY_STATE_TIMEOUT. Further evictions can keep extending it while retries remain eligible. This dates to v1.18.0 (#1542); it's not a regression.
Steps to Reproduce the Problem
- Worker with
WorkerOptions.newBuilder().setStickyQueueScheduleToStartTimeout(Duration.ZERO), so every workflow task replays full history (an eviction or restart during backoff does the same). - Run:
@ActivityInterface public interface Fails { String run(); }
public static class FailsImpl implements Fails {
public String run() { throw new RuntimeException("fail"); }
}
@WorkflowInterface public interface Wf { @WorkflowMethod String run(); }
public static class WfImpl implements Wf {
public String run() {
return Workflow.newLocalActivityStub(Fails.class, LocalActivityOptions.newBuilder()
.setScheduleToCloseTimeout(Duration.ofSeconds(10))
.setLocalRetryThreshold(Duration.ofSeconds(1))
.setRetryOptions(RetryOptions.newBuilder()
.setInitialInterval(Duration.ofSeconds(4))
.setBackoffCoefficient(1)
.setMaximumAttempts(5).build())
.build()).run();
}
}
- Expected:
RETRY_STATE_TIMEOUTafter about 2 attempts, since attempt 3 would start at ~8s with ~2s left, under the 4s backoff. Actual:RETRY_STATE_MAXIMUM_ATTEMPTS_REACHEDafter all 5 attempts.
Specifications
- Version:
mainat 4a4e6b2d (v1.40.0-4) - Platform: macOS, JDK 21, in-process test server
One limitation: the marker value is the first worker's wall clock, so cross-worker clock skew can shorten or lengthen the budget. I have a fix with a regression test ready.
- Ngôn ngữ chính
- Java
- Star
- 433
- Fork
- 257
- Merge trung bình
- 2 ngày 18 giờ
- Pull request đã merge (30 ngày)
- 24
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của temporalio/sdk-java
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
temporalio/sdk-java#3134 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
temporalio/sdk-java#1825 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
temporalio/sdk-java#3132 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
temporalio/sdk-java#3125 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
temporalio/sdk-java#3124 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của temporalio/sdk-java
Issue tương tự
-
waiting-for-triage
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
spring-cloud/spring-cloud-openfeign#1443 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 84/100
ADORSYS-GIS/keycloak-oid4vp-plugin#221 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Upgrade to Spring Pulsar 2.0.8Đang mởstatus: team-only type: dependency-upgrade
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
spring-projects/spring-boot#52099 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
tchiotludo/akhq#3307 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
objectionary/jeo-maven-plugin#1885 ·
Maintainer thường phản hồi trong vòng 4 ngày