compose replay: a bring-up retry orphans the agent's mocks, and the next up can reuse a container that is still stopping
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 38/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- docker, docker-compose, go
- Ambito
- devops, infrastructure, testing-qa
Direzione di ricerca
Start in pkg/client/app/app.go around ComposeDown, the bring-up retry at lines 1457-1468, and the reap barrier near line 694. Trace the straight-line replay setup in pkg/service/replay/replay.go:1285-1488 and mock state in pkg/service/agent/agent.go, including removeStaleComposeAgentWithin. Done means replacement agents retain the session's mocks, compose containers cannot be incorrectly reused after teardown, and a zero-test run cannot pass silently.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Two defects on the docker-compose replay path, found from one CI run. They
compound: the first turns a transient dependency crash into a run that verifies
nothing, and the second is one way to produce that crash. Line numbers are
main (v3.6.67); the behaviour is the same in v3.6.58, where it was
observed.
1. The compose bring-up retry replaces the agent and orphans the session's mocks
pkg/client/app/app.go:1457 retries docker compose up when a dependency
container crashes before the app starts. The retry calls a.ComposeDown()
(:1468), which tears down the injected keploy-agent service along with the
stack — the retry's own comment says so — and then re-issues up.
All replay state lives in the agent process:
- the mock corpus, in
Agent.clientMocks(pkg/service/agent/agent.go:51),
published byfinalizeClientMocks(:752); - the proxy's
MockManager, created byMockOutgoing.
The replacement agent boots empty. Nothing re-arms it: in RunTestSet
(pkg/service/replay/replay.go:1285) the sequence AgentHealthTicker →
MockOutgoing → GetMocks/StoreMocks → SendMockFilterParamsToAgent →
MakeAgentReadyForDockerCompose (:1338–:1488) is straight-line code that has
already run, and AgentHealthTicker returns on the first healthy observation
(pkg/util.go:2790), so the agent going away is never observed.
What the run then produces:
failed to get consumed mocks ... mock manager not found
failed to update mock parameters on agent ... no mocks stored for client ID
app accepted TCP but never completed an HTTP round-trip within the ceiling; firing tests anyway
found no test results for test set with id: test-set-0
The stack itself recovered — the retry's second attempt came up clean — so the
run looked healthy from the outside and reported Total tests: 4, passed 0, failed 0, replay completed successfully, exit 0.
Any transient dependency crash during a compose bring-up reproduces this.
Fix shape (a maintainer's call): after the retry replaces the agent, re-register
the session's mocks with the replacement before any test fires — or fail the
test-set loudly rather than proceeding mockless. Whatever the mechanism, a
zero-test run should not be reachable silently. keploy/keploy#4613 makes such a
run report APP_FAULT instead of PASSED, but it is a reporting fix, not this.
2. The next up can start while the previous stack is still coming down
ComposeDown runs docker compose down --timeout 1 bounded by
composeDownCmdBudget (8s), then a reap barrier that waits on two names —
the agent and the app container (app.go:694). A user service that the down did
not reach, or that is still stopping, is not waited for at all.
The next up then finds it present and reuses it, with its state. In the run
above that container was a zookeeper: every other service printed Creating on
the replay's bring-up, zookeeper printed Starting. It came back with its data
directory, restored the previous broker's session and ephemeral node from its
snapshot, and the new kafka could not claim its own id:
ERROR Error while creating ephemeral at /brokers/ids/1, node already exists and
owner '0x1001bf47ecd0001' does not match current session '0x1001bf56ee30001'
Fatal error during KafkaServer startup
kafka exited 1 and --abort-on-container-exit took the bring-up down — which is
what triggered the retry in defect 1.
I tried the obvious fix (capture the project's containers before the down, wait
for them before the next up) and withdrew it, because review showed it does
not work here and is not safe as written:
- the capture lived on
App, andAgentClient.Setupbuilds a newAppper
call — so at the record→auto-replay boundary, the exact sequence above, the
capture is empty and the guard is a no-op; - with
test.mockingon,composeReusemeans onerun()perApp, so the
capture/consume pair never completes at all; - it waits for a removal nobody issued: when the
downis killed at its budget,
nothing is in flight, so waiting changes nothing — a force-remove has to come
first, asremoveStaleComposeAgentWithinalready does; - and it added a 90s
context.Background()wait on the SIGINT drain path, which
the surrounding code explicitly avoids (app.go:1567).
Recording that here so the next attempt starts past those four traps rather than
into them.
Reproducing
Any compose stack with a service that survives a killed down and carries
identity in its data directory — zookeeper is the clean example — recorded with
keploy record -c "docker compose up" --auto-replay-delay <n>, with a dependency
made to crash once during the replay's bring-up.
- Lingua principale
- Go
- Stelle
- 18.5k
- Fork
- 2.4k
- Merge medio
- 14h 3m
- PR unite (30g)
- 69
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di keploy/keploy
-
[bug]: Postman collection variables containing hyphens are not resolved during importForse già presa Una pull request collegata a questa issue è aperta o già unita. Apertabug keploy
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
I maintainer di solito rispondono entro 1 giorno
-
[bug]: keploy import postman panics on a form-data file fieldForse già presa Una pull request collegata a questa issue è aperta o già unita. Apertakeploy
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
chore: Rename pkg/agent/hooks/linux/comm.go to a more descriptive nameForse già presa @Prateek-og l’ha presa 28 giorni fa. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
keploy/keploy#4566 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
MockReader splits mocks.yaml on indented '---', so keploy cannot read back a file it wrote (replay session fails to start)Forse già presa @Aditya-eddy l’ha presa 43 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
keploy/keploy#4477 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
[bug]: keploy import postman treats Postman folders as API requests and fails with URL is emptyForse già presa @AbhiPra24 l’ha presa 36 giorni fa. Apertabug keploy
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
keploy/keploy#4434 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di keploy/keploy
Issue simili
-
Project submission: 5diveAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
slavakurilyak/awesome-ai-agents#710 ·
I maintainer di solito rispondono entro 1 giorno
-
`renderLinkedIssues` overshoots its byte budget: unresolved and omitted lists are never boundedApertaagent-butler-finding agent-research-recommend bug ready-for-agent
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
jordansmall/spindrift#4614 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
weaviate/weaviate-go-client#485 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
I maintainer di solito rispondono entro 1 giorno
-
Auth server panics in GetProjectById when FindUsersByUID returns an errorForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 1/5 1-3 ore Idoneità per principianti 85/100
litmuschaos/litmus#5641 ·
I maintainer di solito rispondono entro 6 giorni