Local activity scheduleToClose budget restarts when a timer-backed retry runs on replay
Los mantenedores suelen responder en 1 día
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- java
- Área
- backend, distributed-systems
Línea de trabajo
Start in temporal-sdk/src/main/java/io/temporal/internal/statemachines/LocalActivityCallback.java, tracing how firstSkd is parsed and passed into retries. Reproduce with sticky queue scheduling disabled and the listed local activity options; done means replayed retries preserve the original schedule-to-close budget, terminate with RETRY_STATE_TIMEOUT, and a regression test covers the behavior.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Expected Behavior
A local activity that retries through a workflow timer keeps counting ScheduleToCloseTimeout from its first attempt across replay. sdk-core restores the original schedule time from the marker and hands it back with the retry (local_activity_state_machine.rs:147, :646); its proto says it "Must be passed with attempt to the retry LA" (activity_result.proto:97-99). sdk-go sets the schedule time from workflow time (workflow.go:1281) and computes the deadline from it (internal_event_handlers.go:892-894), so both preserve the original scheduling baseline across replay.
Actual Behavior
Java parses firstSkd from the marker into LocalActivityFailedException, then discards it (LocalActivityCallback.java:29-37). The retry uses System.currentTimeMillis() captured when workflow code ran, which on replay is the replay clock. The retry pre-check then sees a budget computed from that clock, so a replay can reset the budget. In the reproduction below the activity runs all 5 attempts and ends with RETRY_STATE_MAXIMUM_ATTEMPTS_REACHED instead of RETRY_STATE_TIMEOUT. Further evictions can keep extending it while retries remain eligible. This dates to v1.18.0 (#1542); it's not a regression.
Steps to Reproduce the Problem
- Worker with
WorkerOptions.newBuilder().setStickyQueueScheduleToStartTimeout(Duration.ZERO), so every workflow task replays full history (an eviction or restart during backoff does the same). - Run:
@ActivityInterface public interface Fails { String run(); }
public static class FailsImpl implements Fails {
public String run() { throw new RuntimeException("fail"); }
}
@WorkflowInterface public interface Wf { @WorkflowMethod String run(); }
public static class WfImpl implements Wf {
public String run() {
return Workflow.newLocalActivityStub(Fails.class, LocalActivityOptions.newBuilder()
.setScheduleToCloseTimeout(Duration.ofSeconds(10))
.setLocalRetryThreshold(Duration.ofSeconds(1))
.setRetryOptions(RetryOptions.newBuilder()
.setInitialInterval(Duration.ofSeconds(4))
.setBackoffCoefficient(1)
.setMaximumAttempts(5).build())
.build()).run();
}
}
- Expected:
RETRY_STATE_TIMEOUTafter about 2 attempts, since attempt 3 would start at ~8s with ~2s left, under the 4s backoff. Actual:RETRY_STATE_MAXIMUM_ATTEMPTS_REACHEDafter all 5 attempts.
Specifications
- Version:
mainat 4a4e6b2d (v1.40.0-4) - Platform: macOS, JDK 21, in-process test server
One limitation: the marker value is the first worker's wall clock, so cross-worker clock skew can shorten or lengthen the budget. I have a fix with a regression test ready.
- Lenguaje dominante
- Java
- Estrellas
- 433
- Forks
- 257
- Merge medio
- 2 d 20 h
- PR fusionados (30 d)
- 20
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de temporalio/sdk-java
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
temporalio/sdk-java#1825 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 38/100
temporalio/sdk-java#3121 ·
Los mantenedores suelen responder en 1 día
-
Allow a timer summary on Workflow.sleep and Workflow.await with timeoutPosiblemente ocupada @sangkyoonnam la tomó hace 4 días. Abiertoenhancement
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
temporalio/sdk-java#3108 ·
Los mantenedores suelen responder en 1 día
-
Warn if the SDK tried to send a payload above a specific size - JavaPosiblemente ocupada @jmaeagle99 la tomó hace 26 días. Abierto
temporalio/sdk-java#3059 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Promise.get(timeout, unit) throws a misleading TimeoutException when the workflow is canceledPosiblemente ocupada @Quinn-With-Two-Ns la tomó hace 41 días. Abierto
temporalio/sdk-java#3026 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de temporalio/sdk-java
Issues similares
-
component/operate kind/feature-request
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
Forge coverage prompts carry text the agent cannot act onPosiblemente ocupada @graalvmbot la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
oracle/graalvm-reachability-metadata#10572 ·
Los mantenedores suelen responder en 1 día
-
[CI] Core CI doesn't run for changes to amoro-format-lance (and amoro-web)Posiblemente ocupada @MarkAlex1234 la tomó hoy. Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
area/docs
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día