internal/run: TestAParkedRootsRunReviewsItsChildBeforeItAnswers fails 100% of the time on one processor
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 68/100
- Issue-Typ
- Bug
- Klarheit
- Größtenteils klar
- Aktivitätsstatus
- Aktiv
- Tech-Stack
- go
- Bereich
- devtools, testing-qa
Rechercherichtung
Run GOMAXPROCS=1 go test -count=1 -run '^TestAParkedRootsRunReviewsItsChildBeforeItAnswers$' ./internal/run/, then inspect internal/run/run.go and internal/run/review_order_test.go, especially Supervisor.Run, releaseStale, passInterval, and parkedRootWorker. Determine whether the supervisor or test scripting causes the ordering failure; done means the pinned and GOMAXPROCS=4 package tests pass without branching on processor count.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
What happened
internal/run went red on two full gate runs, five hours apart, on changes to internal/session. Both times the same test, the same line, the same message, in under half a second:
--- FAIL: TestAParkedRootsRunReviewsItsChildBeforeItAnswers (0.41s)
review_order_test.go:97: outcome = "ran and did not finish", want "done"
The first sighting was at 8b682938d (0.49s) and was cleared with a rerun. The second was at ec9b01cb0 (0.41s), 106 packages ok, one FAIL, load 4.9 at the end of the run.
It was taken for a flake, and it is not one. It fails 100% of the time when the process has a single processor, and has never been seen with two or more.
Measured 2026-09-20 on two heads, 1a692da6e (the base) and ec9b01cb0 (the base plus one change to internal/session), with prebuilt test binaries so nothing came from a build cache, the two arms interleaved run by run so any drift on the box fell on both equally:
| Condition | 1a692da6e |
ec9b01cb0 |
|---|---|---|
| single test, unconstrained | 0/20 | 0/20 |
| whole package, unconstrained | 0/8 | 0/8 |
single test, GOMAXPROCS=1 |
30/30 | 30/30 |
single test, GOMAXPROCS=2 |
0/30 | 0/30 |
single test, GOMAXPROCS=4 |
0/30 | 0/30 |
whole package, GOMAXPROCS=1 |
exit 1, this test only | exit 1, this test only |
-race, whole package, ×2 |
pass, 0 data races | pass, 0 data races |
Every one of the 60 failures is byte-identical to the two gate sightings, between 0.36s and 0.39s. The sibling test in the same file, TestARunningRootsFinishDoesNotCloseTheRunOverAnUnreviewedChild, passes at one processor and at two.
The sub-half-second duration is the part that rules out starvation: the test runs under a 30 second context, so a starved box would show as a timeout near 30 seconds, not as a wrong outcome in 380ms. This is an ordering result.
Replication
Deterministic (no model). One command, no keys, no fixtures, about a second:
GOMAXPROCS=1 go test -count=1 -run '^TestAParkedRootsRunReviewsItsChildBeforeItAnswers$' ./internal/run/
Today that prints the block quoted above and exits 1, every time. The control is the same command with GOMAXPROCS=2, which passes every time.
Field (real models). None. No provider, no key and no model is involved; the test scripts its workers.
Where
internal/run/run.go. Supervisor.Run is the loop that answers the outcome, and OutcomeIncomplete is the constant whose text is ran and did not finish. The scenario is in internal/run/review_order_test.go: the supervisor is built with two slots and the test forces an ordering with childDone and lateTried, with no sleeps, so a root parks on its child while a round its worker had already started still ends its tool.
Searchable strings: OutcomeIncomplete, func (s *Supervisor) Run, releaseStale, passInterval, parkedRootWorker.
What has not been established: which line assumes a second processor, and therefore whether the defect is in the supervisor or in the test's own scripting. The scheduling story below is consistent with everything measured and is not proof of a mechanism.
The fix
A run's answer must not depend on how many processors the process was given. Supervisor.Run should answer done for this scripted order on one processor exactly as it does on two, or, if the test is asserting an order the supervisor never promised, the test should assert the promise it actually makes and say so.
Whichever it is, the outcome is decided by the same constants today and no person-facing wording changes, so no manual page is owed by this fix unless the outcome a run reports changes, which it should not.
Acceptance
- e2e:
GOMAXPROCS=1 go test -count=1 ./internal/run/exits 0, with no--- FAILline forTestAParkedRootsRunReviewsItsChildBeforeItAnswers. This is the command that fails today. - e2e: the control,
GOMAXPROCS=4 go test -count=1 ./internal/run/, still exits 0, so the fix does not trade one parallelism for another. - Unit: the scenario pins its own parallelism rather than inheriting the runner's — the test sets
runtime.GOMAXPROCS(1)for its own duration and restores it — so the guarantee is that the run answersdoneon one processor, asserted in the test itself and not dependent on how the gate happens to be configured. Without this the fix is provable only by remembering to pass an environment variable, which is a regression test that can silently stop checking. - And the fix does not read the processor count. Nothing in
internal/runmay branch onruntime.GOMAXPROCSorruntime.NumCPUto satisfy the line above. That branch is the cheapest way to make a pinned scenario pass while leaving the ordering assumption in place for every machine that behaves like one processor under load — which, on the reading above, is how this reached two gates. The shape that counts: the pinned scenario passes, the unconstrained runs keep passing, and the diff contains no read of the processor count.
Two open questions, neither answered here
What parallelism does the gate run under? If a full check runs unconstrained on a many-core box, then two sightings in roughly a dozen full runs is not explained by this result, and the story that load transiently narrows effective parallelism is doing a lot of work. If the gate constrains GOMAXPROCS, the explanation is straightforward. This was not checked.
A lock reading on the test box does not mean the box is quiet. While these runs were taken, two suites were running from other rigs (go test ./internal/tui3 and go test ./internal/session) and the suite flock read FREE throughout, because runs inside cell sandboxes do not take it. Load stayed between 3.2 and 5.6 for the whole experiment. That does not affect this result — 60/60 against 0/120 is not a load effect — but a reader should not take "the lock was free" as "nothing else was running" on this box.
- Vorherrschende Sprache
- Go
- Sterne
- 115
- Forks
- 14
- Ø Merge
- 9 Std. 37 Min.
- Gemergte PRs (30 T.)
- 755
Entwicklungsumgebung
Die Einrichtungsdateien dieses Projekts haben wir noch nicht geprüft. Beginnen Sie mit der README; die allgemeinen Schritte stehen in unserem Leitfaden für den ersten Beitrag.
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus Agent-Field/CodeAF
-
area:chat bug sev:papercut
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 86/100
Agent-Field/CodeAF#1592 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:headless bug sev:critical
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Agent-Field/CodeAF#1566 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:chat feature
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Agent-Field/CodeAF#1510 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:tests bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
Agent-Field/CodeAF#1489 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:chat bug good first issue sev:papercut
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Agent-Field/CodeAF#1470 ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in Agent-Field/CodeAF
Ähnliche Issues
-
Remove CAAPFOffenkind/chore kind/cleanup needs-area
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 86/100
rancher/turtles#2848 · 3 Kommentare ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Maintainer antworten meist innerhalb von 1 Tag
-
good first issue
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
Maintainer antworten meist innerhalb von 1 Tag
-
priority: low 🌱 type: enhancement 💅🏼
Schwierigkeit 2/5 Ein halber Tag Anfängerfreundlichkeit 84/100
nebari-dev/llm-serving-pack#199 ·
Maintainer antworten meist innerhalb von 3 Tagen
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 90/100
kedacore/keda#8225 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag