test server: retrying activity never times out by ScheduleToClose when its expiration falls on a whole second (`getNanos() != 0` guard in `TestServiceRetryState`)
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 2/5
- Geschätzter Aufwand
- 1-3 Stunden
- Anfängerfreundlichkeit
- 78/100
Rechercherichtung
Lies TestServiceRetryState.getBackoffIntervalInSeconds im Java-Testserver, insbesondere die im Issue beschriebene Ablaufprüfung. Vergleiche den Zeitpunkt des nächsten Versuchs mit dem Ablaufzeitpunkt, auch wenn dessen Nanosekunden null sind, und führe anschließend die relevanten Retry-State-Tests aus. Fertig ist es, wenn ein Activity bei ScheduleToClose ein Timeout hat, unabhängig davon, ob der Ablauf auf eine volle Sekunde fällt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Summary
In the Java test server (temporal-test-server), an activity that keeps failing with a retryable error is supposed to stop retrying once the next attempt would start after its ScheduleToClose deadline. In TestServiceRetryState.getBackoffIntervalInSeconds this check is guarded by expirationTime.getNanos() != 0:
if (expirationTime.getNanos() != 0
&& Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}
The test server's clock has millisecond precision, so the expiration timestamp has nanos == 0 whenever the activity was scheduled exactly on a whole second (≈1 in 1000 schedules). In that case the check is skipped, and nothing else closes the activity: the ScheduleToClose timer is bound to the attempt number at scheduling time and is discarded as outdated after the first retry. The activity then retries forever (we observed 362 attempts at 20 s intervals well past a 2 h deadline), and the workflow never completes.
The guard looks unnecessary: the constructor already maps "no expiration" to Timestamps.MAX_VALUE, so comparing against it is safe.
Reproduction
Minimal workflow (Kotlin, SDK 1.25.1, TestWorkflowEnvironment with time skipping):
- activity always throws a retryable
ApplicationFailure; ScheduleToCloseTimeout = 2h,RetryOptions(initialInterval = 20m, backoffCoefficient = 1.0), nomaximumAttempts;- workflow catches
ActivityFailureand returns.
Running many independent executions and waiting 5 s (wall clock) for each result:
| test server | executions | never completed | ActivityTaskScheduled.eventTime.nanos of the stuck ones |
|---|---|---|---|
| 1.25.1 | 4514 | 4 | 0 in all 4 |
| 1.40.0 | 409 | 1 | 0 |
1.25.1 with the getNanos() != 0 && part removed |
12000 | 0 (11 schedules had nanos = 0) | — |
A deterministic reproduction is possible by freezing the test server clock on a whole second (we did it with a test-scoped TestServicesStarter variant using Clock.fixed(...)): without the change the activity never closes; with it, it closes by RETRY_STATE_TIMEOUT.
Expected
RETRY_STATE_TIMEOUT whenever the next attempt would be scheduled after the expiration time, regardless of the sub-second part of the expiration timestamp.
Suggested fix
if (Timestamps.compare(nextScheduleTime, expirationTime) > 0) {
return new BackoffInterval(RetryState.RETRY_STATE_TIMEOUT);
}
Versions checked: 1.25.1, 1.26.0, 1.28.0, 1.30.0, 1.34.0, 1.36.0, 1.40.0 — the same condition is present. We currently work around it by shadowing the class in our test classpath.
- Vorherrschende Sprache
- Java
- Sterne
- 433
- Forks
- 257
- Ø Merge
- 2 T. 18 Std.
- Gemergte PRs (30 T.)
- 24
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Keine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus temporalio/sdk-java
-
enhancement
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
temporalio/sdk-java#1825 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 55/100
temporalio/sdk-java#3132 ·
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 54/100
temporalio/sdk-java#3125 · 2 Kommentare ·
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 55/100
temporalio/sdk-java#3124 ·
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
temporalio/sdk-java#3122 ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in temporalio/sdk-java
Ähnliche Issues
-
backend
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
bcgov/nr-forest-client#2524 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 67/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 74/100
Maintainer antworten meist innerhalb von 1 Tag
-
team:Lumberjack
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
OpenLiberty/open-liberty#35998 ·
Maintainer antworten meist innerhalb von 1 Tag
-
[BUG] SQS SendMessageBatch accepts more than 10 entries instead of TooManyEntriesInBatchRequestEvtl. vergeben Ein verknüpfter Pull Request ist offen oder bereits gemergt. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 67/100
floci-io/floci#5319 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag