Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

test server: retrying activity never times out by ScheduleToClose when its expiration falls on a whole second (`getNanos() != 0` guard in `TestServiceRetryState`)

未关闭 适合新手
#3,134 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
78/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
java
领域
backend

调研方向

阅读 Java 测试服务器中的 TestServiceRetryState.getBackoffIntervalInSeconds,尤其是 issue 中描述的过期检查。即使纳秒为零,也要比较下一次尝试时间和过期时间,然后运行相关的重试状态测试。完成标准是:无论过期时间是否落在整秒上,活动都会在 ScheduleToClose 时超时。

由索引模型根据 Issue 内容生成。

描述

Summary

In the Java test server (temporal-test-server), an activity that keeps failing with a retryable error is supposed to stop retrying once the next attempt would start after its ScheduleToClose deadline. In TestServiceRetryState.getBackoffIntervalInSeconds this check is guarded by expirationTime.getNanos() != 0:

if (expirationTime.getNanos() != 0
    && Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
  return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}

The test server's clock has millisecond precision, so the expiration timestamp has nanos == 0 whenever the activity was scheduled exactly on a whole second (≈1 in 1000 schedules). In that case the check is skipped, and nothing else closes the activity: the ScheduleToClose timer is bound to the attempt number at scheduling time and is discarded as outdated after the first retry. The activity then retries forever (we observed 362 attempts at 20 s intervals well past a 2 h deadline), and the workflow never completes.

The guard looks unnecessary: the constructor already maps "no expiration" to Timestamps.MAX_VALUE, so comparing against it is safe.

Reproduction

Minimal workflow (Kotlin, SDK 1.25.1, TestWorkflowEnvironment with time skipping):

  • activity always throws a retryable ApplicationFailure;
  • ScheduleToCloseTimeout = 2h, RetryOptions(initialInterval = 20m, backoffCoefficient = 1.0), no maximumAttempts;
  • workflow catches ActivityFailure and returns.

Running many independent executions and waiting 5 s (wall clock) for each result:

test server executions never completed ActivityTaskScheduled.eventTime.nanos of the stuck ones
1.25.1 4514 4 0 in all 4
1.40.0 409 1 0
1.25.1 with the getNanos() != 0 && part removed 12000 0 (11 schedules had nanos = 0) —

A deterministic reproduction is possible by freezing the test server clock on a whole second (we did it with a test-scoped TestServicesStarter variant using Clock.fixed(...)): without the change the activity never closes; with it, it closes by RETRY_STATE_TIMEOUT.

Expected

RETRY_STATE_TIMEOUT whenever the next attempt would be scheduled after the expiration time, regardless of the sub-second part of the expiration timestamp.

Suggested fix

if (Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
  return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}

Versions checked: 1.25.1, 1.26.0, 1.28.0, 1.30.0, 1.34.0, 1.36.0, 1.40.0 — the same condition is present. We currently work around it by shadowing the class in our test classpath.

主要语言
Java
星标
433
派生
257
平均合并
2 天 18 小时
30 天内合并 PR
24

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

temporalio/sdk-java 的其他 Issue

查看 temporalio/sdk-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。