Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

compose replay: a bring-up retry orphans the agent's mocks, and the next up can reuse a container that is still stopping

Aperta
#4,614 6 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
38/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
docker, docker-compose, go

Direzione di ricerca

Start in pkg/client/app/app.go around ComposeDown, the bring-up retry at lines 1457-1468, and the reap barrier near line 694. Trace the straight-line replay setup in pkg/service/replay/replay.go:1285-1488 and mock state in pkg/service/agent/agent.go, including removeStaleComposeAgentWithin. Done means replacement agents retain the session's mocks, compose containers cannot be incorrectly reused after teardown, and a zero-test run cannot pass silently.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Two defects on the docker-compose replay path, found from one CI run. They
compound: the first turns a transient dependency crash into a run that verifies
nothing, and the second is one way to produce that crash. Line numbers are
main (v3.6.67); the behaviour is the same in v3.6.58, where it was
observed.

1. The compose bring-up retry replaces the agent and orphans the session's mocks

pkg/client/app/app.go:1457 retries docker compose up when a dependency
container crashes before the app starts. The retry calls a.ComposeDown()
(:1468), which tears down the injected keploy-agent service along with the
stack
— the retry's own comment says so — and then re-issues up.

All replay state lives in the agent process:

  • the mock corpus, in Agent.clientMocks (pkg/service/agent/agent.go:51),
    published by finalizeClientMocks (:752);
  • the proxy's MockManager, created by MockOutgoing.

The replacement agent boots empty. Nothing re-arms it: in RunTestSet
(pkg/service/replay/replay.go:1285) the sequence AgentHealthTicker →
MockOutgoing → GetMocks/StoreMocks → SendMockFilterParamsToAgent →
MakeAgentReadyForDockerCompose (:1338–:1488) is straight-line code that has
already run, and AgentHealthTicker returns on the first healthy observation
(pkg/util.go:2790), so the agent going away is never observed.

What the run then produces:

failed to get consumed mocks ... mock manager not found
failed to update mock parameters on agent ... no mocks stored for client ID
app accepted TCP but never completed an HTTP round-trip within the ceiling; firing tests anyway
found no test results for test set with id: test-set-0

The stack itself recovered — the retry's second attempt came up clean — so the
run looked healthy from the outside and reported Total tests: 4, passed 0, failed 0, replay completed successfully, exit 0.

Any transient dependency crash during a compose bring-up reproduces this.

Fix shape (a maintainer's call): after the retry replaces the agent, re-register
the session's mocks with the replacement before any test fires — or fail the
test-set loudly rather than proceeding mockless. Whatever the mechanism, a
zero-test run should not be reachable silently. keploy/keploy#4613 makes such a
run report APP_FAULT instead of PASSED, but it is a reporting fix, not this.

2. The next up can start while the previous stack is still coming down

ComposeDown runs docker compose down --timeout 1 bounded by
composeDownCmdBudget (8s), then a reap barrier that waits on two names —
the agent and the app container (app.go:694). A user service that the down did
not reach, or that is still stopping, is not waited for at all.

The next up then finds it present and reuses it, with its state. In the run
above that container was a zookeeper: every other service printed Creating on
the replay's bring-up, zookeeper printed Starting. It came back with its data
directory, restored the previous broker's session and ephemeral node from its
snapshot, and the new kafka could not claim its own id:

ERROR Error while creating ephemeral at /brokers/ids/1, node already exists and
owner '0x1001bf47ecd0001' does not match current session '0x1001bf56ee30001'
Fatal error during KafkaServer startup

kafka exited 1 and --abort-on-container-exit took the bring-up down — which is
what triggered the retry in defect 1.

I tried the obvious fix (capture the project's containers before the down, wait
for them before the next up) and withdrew it, because review showed it does
not work here and is not safe as written:

  • the capture lived on App, and AgentClient.Setup builds a new App per
    call — so at the record→auto-replay boundary, the exact sequence above, the
    capture is empty and the guard is a no-op;
  • with test.mocking on, composeReuse means one run() per App, so the
    capture/consume pair never completes at all;
  • it waits for a removal nobody issued: when the down is killed at its budget,
    nothing is in flight, so waiting changes nothing — a force-remove has to come
    first, as removeStaleComposeAgentWithin already does;
  • and it added a 90s context.Background() wait on the SIGINT drain path, which
    the surrounding code explicitly avoids (app.go:1567).

Recording that here so the next attempt starts past those four traps rather than
into them.

Reproducing

Any compose stack with a service that survives a killed down and carries
identity in its data directory — zookeeper is the clean example — recorded with
keploy record -c "docker compose up" --auto-replay-delay <n>, with a dependency
made to crash once during the replay's bring-up.

Lingua principale
Go
Stelle
18.5k
Fork
2.4k
Merge medio
14h 3m
PR unite (30g)
69

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di keploy/keploy

Tutte le issue di keploy/keploy

Issue simili

Altre issue su Go

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.