Flaky on JDK 11: BackgroundDrainerMidDrainCapabilityGapTest.testDeliveringBetweenTwoGapWindowsGrantsAFreshSettleBudget
Los mantenedores suelen responder en 3 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 42/100
- Tipo de issue
- Error
- Claridad
- Necesita aclaración
- Estado de actividad
- Activo
- Stack tecnológico
- java
- Área
- testing-qa
Línea de trabajo
Reproduce BackgroundDrainerMidDrainCapabilityGapTest#testDeliveringBetweenTwoGapWindowsGrantsAFreshSettleBudget con el comando Maven proporcionado en JDK 11. Inspecciona BackgroundDrainer.java alrededor de las líneas 525-527, 563-565, 618 y 683-685, y luego captura la línea de abandono que falla para distinguir una condición de carrera del harness de TestWebSocketServer de un fallo al restablecer el episodio. Se considera terminado cuando el test deja de fallar de forma intermitente y se cubre el comportamiento correcto de restablecimiento o de reconocimiento determinista.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
BackgroundDrainerMidDrainCapabilityGapTest#testDeliveringBetweenTwoGapWindowsGrantsAFreshSettleBudget fails intermittently on JDK 11, on main as well as on feature branches. Found while reviewing #99; it is not caused by that PR.
Symptom
java.lang.AssertionError: delivering between the windows ends the episode, so neither window
reaches the threshold and the slot must still drain [attempts=20] expected:<SUCCESS> but was:<FAILED>
at ...BackgroundDrainerMidDrainCapabilityGapTest.lambda$testDeliveringBetweenTwoGapWindowsGrantsAFreshSettleBudget$0(BackgroundDrainerMidDrainCapabilityGapTest.java:240)
at ...TestUtils.assertMemoryLeak(TestUtils.java:135)
at ...BackgroundDrainerMidDrainCapabilityGapTest.testDeliveringBetweenTwoGapWindowsGrantsAFreshSettleBudget(BackgroundDrainerMidDrainCapabilityGapTest.java:201)
The drainer ends the episode with DrainOutcome.FAILED and quarantines the slot, where the test expects SUCCESS. factory.attempts() is 20 — the exact end of the second scripted gap window.
Reproduction
JAVA_HOME=~/.sdkman/candidates/java/11.0.26-amzn \
mvn -pl core test -Dtest='BackgroundDrainerMidDrainCapabilityGapTest'
Re-run several times — it does not fail every run.
Evidence that it is flaky and not revision-specific
JDK 11.0.26-amzn (Corretto), Linux x86-64, repeated runs of the same test:
| revision | failures / runs |
|---|---|
981bdb0 (main) |
4 / 7 |
065c7be (PR #99 branch) |
3 / 7 |
Both revisions both pass and fail it. It initially looked like a regression on the PR branch (branch failed, main passed), but the outcome also inverts depending on how the test is invoked: running the single method in isolation vs. running the whole class flipped which revision failed, in both directions, across repeats. That rules out a code difference between the two revisions as the cause.
Not observed on JDK 26: a full mvn -pl core test there passed (3456 tests, 0 failures) — though that is a single observation, not a claim that the JDK matters.
What I ruled out
- Machine load. Reproduced on an idle machine (load ~1.6), not only under parallel builds.
- The 60 s wall-clock settle budget.
RECONNECT_MAX_DURATION_MILLIS = 60_000in the test, backoff 1–4 ms, and the whole run finishes in ~10 s, socapabilityGapElapsedNanos >= reconnectBudgetNanos(BackgroundDrainer.java:618) is not the terminal that fires. - The settle-budget accumulators failing to reset as a group.
capabilityGapAttempts,capabilityGapElapsedNanosandlastCapabilityGapNanosare reset together at all three sites (BackgroundDrainer.java:525-527, 563-565, 683-685), and a captured log of a passing run shows the episode restarting correctly — the second gap window begins again atattempt 1:
21:11:51.088 ... durable-ack capability gap mid-drain (...), re-entering settle budget
21:11:51.089 ... attempt 1: durable-ack unavailable, retrying after backoff
... attempts 2..8 ...
21:11:56.044 ... durable-ack capability gap mid-drain (...), re-entering settle budget <- ~5 s delivering gap
21:11:56.054 ... attempt 1: durable-ack unavailable, retrying after backoff
What I did not determine
Root cause. Two candidates remain, and I could not separate them:
- Test-harness race. The scripted delivering session (connection 2 acks frames 0 and 1, then drops —
drops.put(2, 1L)) sometimes loses the race and never advances the watermark, so no reset happens, the two 9-sweep windows accumulate to 18 ≥DEFAULT_MAX_DURABLE_ACK_MISMATCH_ATTEMPTS(16), andFAILEDis the correct outcome for what actually happened on the wire. - A genuine intermittent failure of the episode reset in
BackgroundDrainer.
The test drives a real TestWebSocketServer over a real loopback socket, so the delivering session's ack is genuinely timing-dependent.
The decisive diagnostic is the give-up line from a failing run, which reports which terminal fired and the elapsed time:
drainer giving up on slot {} after {} durable-ack-mismatch attempts ({}ms): {}
(BackgroundDrainer.java:620). If it reports ~18 attempts, candidate 1 is the explanation and the fix belongs in the test harness (make the delivering session's ack deterministic before the connection drops). I only managed to capture that log on passing runs.
Impact
Low for CI — build-jdk8 is the only job that runs tests, so this never reaches a pipeline today. It costs developer time on local JDK 11+ runs, and it is masking whichever of the two candidates above is true.
🤖 Generated with Claude Code
- Lenguaje dominante
- Java
- Estrellas
- 10
- Forks
- 6
- Merge medio
- 8 d 7 h
- PR fusionados (30 d)
- 3
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de questdb/java-questdb-client
-
enhancement
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
questdb/java-questdb-client#83 · 1 comentario ·
Los mantenedores suelen responder en 3 días
-
1 million of district symbols limitPosiblemente ocupada @jovfer la tomó hace 49 días. Abiertobug
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
questdb/java-questdb-client#80 ·
Los mantenedores suelen responder en 3 días
-
InternalError thrown from MmapSegmentRecoveryFaultTest.testScanFaultOnMapPastEofIsHandledAnyFilesystem()Quizá libre de nuevo Un pull request para esta issue se cerró sin fusionarse. Abiertobug
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
questdb/java-questdb-client#69 ·
Los mantenedores suelen responder en 3 días
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 58/100
questdb/java-questdb-client#18 · 1 comentario ·
Los mantenedores suelen responder en 3 días
Todos los issues de questdb/java-questdb-client
Issues similares
-
waiting-for-triage
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
spring-cloud/spring-cloud-openfeign#1443 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 1-3 horas Aptitud para principiantes 84/100
ADORSYS-GIS/keycloak-oid4vp-plugin#221 ·
Los mantenedores suelen responder en 2 días
-
Upgrade to Spring Pulsar 2.0.8Abiertostatus: team-only type: dependency-upgrade
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
spring-projects/spring-boot#52099 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 67/100
tchiotludo/akhq#3307 · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
objectionary/jeo-maven-plugin#1885 ·
Los mantenedores suelen responder en 4 días