test server: retrying activity never times out by ScheduleToClose when its expiration falls on a whole second (`getNanos() != 0` guard in `TestServiceRetryState`)
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 2/5
- Tempo estimado
- 1-3 horas
- Facilidade para iniciantes
- 78/100
Direção de pesquisa
Leia TestServiceRetryState.getBackoffIntervalInSeconds no servidor de teste Java, especialmente a verificação de expiração descrita na issue. Compare o horário da próxima tentativa com o horário de expiração, mesmo quando os nanossegundos forem zero, e depois execute os testes relevantes do estado de repetição. O trabalho estará concluído quando uma atividade atingir o timeout em ScheduleToClose, independentemente de a expiração cair em um segundo exato.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Summary
In the Java test server (temporal-test-server), an activity that keeps failing with a retryable error is supposed to stop retrying once the next attempt would start after its ScheduleToClose deadline. In TestServiceRetryState.getBackoffIntervalInSeconds this check is guarded by expirationTime.getNanos() != 0:
if (expirationTime.getNanos() != 0
&& Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}
The test server's clock has millisecond precision, so the expiration timestamp has nanos == 0 whenever the activity was scheduled exactly on a whole second (≈1 in 1000 schedules). In that case the check is skipped, and nothing else closes the activity: the ScheduleToClose timer is bound to the attempt number at scheduling time and is discarded as outdated after the first retry. The activity then retries forever (we observed 362 attempts at 20 s intervals well past a 2 h deadline), and the workflow never completes.
The guard looks unnecessary: the constructor already maps "no expiration" to Timestamps.MAX_VALUE, so comparing against it is safe.
Reproduction
Minimal workflow (Kotlin, SDK 1.25.1, TestWorkflowEnvironment with time skipping):
- activity always throws a retryable
ApplicationFailure; ScheduleToCloseTimeout = 2h,RetryOptions(initialInterval = 20m, backoffCoefficient = 1.0), nomaximumAttempts;- workflow catches
ActivityFailureand returns.
Running many independent executions and waiting 5 s (wall clock) for each result:
| test server | executions | never completed | ActivityTaskScheduled.eventTime.nanos of the stuck ones |
|---|---|---|---|
| 1.25.1 | 4514 | 4 | 0 in all 4 |
| 1.40.0 | 409 | 1 | 0 |
1.25.1 with the getNanos() != 0 && part removed |
12000 | 0 (11 schedules had nanos = 0) | — |
A deterministic reproduction is possible by freezing the test server clock on a whole second (we did it with a test-scoped TestServicesStarter variant using Clock.fixed(...)): without the change the activity never closes; with it, it closes by RETRY_STATE_TIMEOUT.
Expected
RETRY_STATE_TIMEOUT whenever the next attempt would be scheduled after the expiration time, regardless of the sub-second part of the expiration timestamp.
Suggested fix
if (Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}
Versions checked: 1.25.1, 1.26.0, 1.28.0, 1.30.0, 1.34.0, 1.36.0, 1.40.0 — the same condition is present. We currently work around it by shadowing the class in our test classpath.
- Linguagem predominante
- Java
- Estrelas
- 433
- Forks
- 257
- Merge médio
- 2d 18h
- PRs com merge (30d)
- 24
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Sem modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de temporalio/sdk-java
-
enhancement
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
temporalio/sdk-java#1825 ·
Mantenedores costumam responder em até 1 dia
-
Proposal: handle SIGTERM by default to initiate graceful worker shutdownTalvez já em andamento @eamsden assumiu hoje. Aberta
temporalio/sdk-java#3135 · 1 comentário · 1 responsável ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 55/100
temporalio/sdk-java#3132 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
enhancement
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 54/100
temporalio/sdk-java#3125 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
enhancement
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 55/100
temporalio/sdk-java#3124 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de temporalio/sdk-java
Issues semelhantes
-
Fix Math.ceilDiv wrong result for exact positive divisionsTalvez já em andamento @pamod-madubashana assumiu hoje. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
scala-native/scala-native#5094 ·
Mantenedores costumam responder em até 1 dia
-
[Bug] AI unread message badge counts a batch of new bubbles as one messageTalvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 74/100
apache/rocketmq-dashboard#5784 ·
Mantenedores costumam responder em até 3 dias
-
[i18n] 安装实例完成后的成功提示未正确本地化Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 62/100
PCL-Community/PCL-CE#3658 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 66/100
apache/skywalking#14127 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 82/100
Mantenedores costumam responder em até 1 dia