Local activity scheduleToClose budget restarts when a timer-backed retry runs on replay
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- java
- Domain
- backend, distributed-systems
Research direction
Start in temporal-sdk/src/main/java/io/temporal/internal/statemachines/LocalActivityCallback.java, tracing how firstSkd is parsed and passed into retries. Reproduce with sticky queue scheduling disabled and the listed local activity options; done means replayed retries preserve the original schedule-to-close budget, terminate with RETRY_STATE_TIMEOUT, and a regression test covers the behavior.
Written by the indexing model from the issue text.
Description
Expected Behavior
A local activity that retries through a workflow timer keeps counting ScheduleToCloseTimeout from its first attempt across replay. sdk-core restores the original schedule time from the marker and hands it back with the retry (local_activity_state_machine.rs:147, :646); its proto says it "Must be passed with attempt to the retry LA" (activity_result.proto:97-99). sdk-go sets the schedule time from workflow time (workflow.go:1281) and computes the deadline from it (internal_event_handlers.go:892-894), so both preserve the original scheduling baseline across replay.
Actual Behavior
Java parses firstSkd from the marker into LocalActivityFailedException, then discards it (LocalActivityCallback.java:29-37). The retry uses System.currentTimeMillis() captured when workflow code ran, which on replay is the replay clock. The retry pre-check then sees a budget computed from that clock, so a replay can reset the budget. In the reproduction below the activity runs all 5 attempts and ends with RETRY_STATE_MAXIMUM_ATTEMPTS_REACHED instead of RETRY_STATE_TIMEOUT. Further evictions can keep extending it while retries remain eligible. This dates to v1.18.0 (#1542); it's not a regression.
Steps to Reproduce the Problem
- Worker with
WorkerOptions.newBuilder().setStickyQueueScheduleToStartTimeout(Duration.ZERO), so every workflow task replays full history (an eviction or restart during backoff does the same). - Run:
@ActivityInterface public interface Fails { String run(); }
public static class FailsImpl implements Fails {
public String run() { throw new RuntimeException("fail"); }
}
@WorkflowInterface public interface Wf { @WorkflowMethod String run(); }
public static class WfImpl implements Wf {
public String run() {
return Workflow.newLocalActivityStub(Fails.class, LocalActivityOptions.newBuilder()
.setScheduleToCloseTimeout(Duration.ofSeconds(10))
.setLocalRetryThreshold(Duration.ofSeconds(1))
.setRetryOptions(RetryOptions.newBuilder()
.setInitialInterval(Duration.ofSeconds(4))
.setBackoffCoefficient(1)
.setMaximumAttempts(5).build())
.build()).run();
}
}
- Expected:
RETRY_STATE_TIMEOUTafter about 2 attempts, since attempt 3 would start at ~8s with ~2s left, under the 4s backoff. Actual:RETRY_STATE_MAXIMUM_ATTEMPTS_REACHEDafter all 5 attempts.
Specifications
- Version:
mainat 4a4e6b2d (v1.40.0-4) - Platform: macOS, JDK 21, in-process test server
One limitation: the marker value is the first worker's wall clock, so cross-worker clock skew can shorten or lengthen the budget. I have a fix with a regression test ready.
- Dominant language
- Java
- Stars
- 433
- Forks
- 257
- Avg merge
- 2d 22h
- Merged PRs (30d)
- 19
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from temporalio/sdk-java
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
temporalio/sdk-java#2676 · 8 comments · 2 reactions ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
temporalio/sdk-java#1825 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 4/5 3-5 days Newbie friendliness 55/100
temporalio/sdk-java#3108 ·
Maintainers usually reply within 1 day
-
Warn if the SDK tried to send a payload above a specific size - JavaPossibly taken @jmaeagle99 claimed this 24 days ago. Open
temporalio/sdk-java#3059 · 1 assignee ·
Maintainers usually reply within 1 day
-
Promise.get(timeout, unit) throws a misleading TimeoutException when the workflow is canceledMay be free again @Quinn-With-Two-Ns claimed this 39 days ago, and no pull request is open. Open
temporalio/sdk-java#3026 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
All issues in temporalio/sdk-java
Similar issues
-
type: possible bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
grimmory-tools/grimmory#2850 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
aoqia194/leaf-loader#19 ·