test server: retrying activity never times out by ScheduleToClose when its expiration falls on a whole second (`getNanos() != 0` guard in `TestServiceRetryState`)
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
阅读 Java 测试服务器中的 TestServiceRetryState.getBackoffIntervalInSeconds,尤其是 issue 中描述的过期检查。即使纳秒为零,也要比较下一次尝试时间和过期时间,然后运行相关的重试状态测试。完成标准是:无论过期时间是否落在整秒上,活动都会在 ScheduleToClose 时超时。
由索引模型根据 Issue 内容生成。
描述
Summary
In the Java test server (temporal-test-server), an activity that keeps failing with a retryable error is supposed to stop retrying once the next attempt would start after its ScheduleToClose deadline. In TestServiceRetryState.getBackoffIntervalInSeconds this check is guarded by expirationTime.getNanos() != 0:
if (expirationTime.getNanos() != 0
&& Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}
The test server's clock has millisecond precision, so the expiration timestamp has nanos == 0 whenever the activity was scheduled exactly on a whole second (≈1 in 1000 schedules). In that case the check is skipped, and nothing else closes the activity: the ScheduleToClose timer is bound to the attempt number at scheduling time and is discarded as outdated after the first retry. The activity then retries forever (we observed 362 attempts at 20 s intervals well past a 2 h deadline), and the workflow never completes.
The guard looks unnecessary: the constructor already maps "no expiration" to Timestamps.MAX_VALUE, so comparing against it is safe.
Reproduction
Minimal workflow (Kotlin, SDK 1.25.1, TestWorkflowEnvironment with time skipping):
- activity always throws a retryable
ApplicationFailure; ScheduleToCloseTimeout = 2h,RetryOptions(initialInterval = 20m, backoffCoefficient = 1.0), nomaximumAttempts;- workflow catches
ActivityFailureand returns.
Running many independent executions and waiting 5 s (wall clock) for each result:
| test server | executions | never completed | ActivityTaskScheduled.eventTime.nanos of the stuck ones |
|---|---|---|---|
| 1.25.1 | 4514 | 4 | 0 in all 4 |
| 1.40.0 | 409 | 1 | 0 |
1.25.1 with the getNanos() != 0 && part removed |
12000 | 0 (11 schedules had nanos = 0) | — |
A deterministic reproduction is possible by freezing the test server clock on a whole second (we did it with a test-scoped TestServicesStarter variant using Clock.fixed(...)): without the change the activity never closes; with it, it closes by RETRY_STATE_TIMEOUT.
Expected
RETRY_STATE_TIMEOUT whenever the next attempt would be scheduled after the expiration time, regardless of the sub-second part of the expiration timestamp.
Suggested fix
if (Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}
Versions checked: 1.25.1, 1.26.0, 1.28.0, 1.30.0, 1.34.0, 1.36.0, 1.40.0 — the same condition is present. We currently work around it by shadowing the class in our test classpath.
- 主要语言
- Java
- 星标
- 433
- 派生
- 257
- 平均合并
- 2 天 18 小时
- 30 天内合并 PR
- 24
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
temporalio/sdk-java 的其他 Issue
-
enhancement
难度 2/5 1-3 小时 新手友好度 65/100
temporalio/sdk-java#1825 ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 55/100
temporalio/sdk-java#3132 ·
维护者通常 1 天内回复
-
enhancement
难度 4/5 3-5 天 新手友好度 54/100
temporalio/sdk-java#3125 · 2 条评论 ·
维护者通常 1 天内回复
-
enhancement
难度 3/5 1-2 天 新手友好度 55/100
temporalio/sdk-java#3124 ·
维护者通常 1 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 35/100
temporalio/sdk-java#3122 ·
维护者通常 1 天内回复
查看 temporalio/sdk-java 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 72/100
NationalSecurityAgency/ghidra#9748 ·
维护者通常 1 天内回复
-
enhancement
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
spring-mcp-tools
难度 2/5 1-3 小时 新手友好度 85/100
explyt/spring-plugin#591 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
jenkinsci/build-monitor-plugin#1367 ·
维护者通常 1 天内回复
-
waiting-for-triage
难度 1/5 1 小时以内 新手友好度 72/100
spring-cloud/spring-cloud-openfeign#1443 ·
维护者通常 1 天内回复