internal/run: TestAParkedRootsRunReviewsItsChildBeforeItAnswers fails 100% of the time on one processor
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 68/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- go
- 領域
- devtools, testing-qa
調査の方向性
Run GOMAXPROCS=1 go test -count=1 -run '^TestAParkedRootsRunReviewsItsChildBeforeItAnswers$' ./internal/run/, then inspect internal/run/run.go and internal/run/review_order_test.go, especially Supervisor.Run, releaseStale, passInterval, and parkedRootWorker. Determine whether the supervisor or test scripting causes the ordering failure; done means the pinned and GOMAXPROCS=4 package tests pass without branching on processor count.
索引モデルが issue の本文から書いたものです。
説明
What happened
internal/run went red on two full gate runs, five hours apart, on changes to internal/session. Both times the same test, the same line, the same message, in under half a second:
--- FAIL: TestAParkedRootsRunReviewsItsChildBeforeItAnswers (0.41s)
review_order_test.go:97: outcome = "ran and did not finish", want "done"
The first sighting was at 8b682938d (0.49s) and was cleared with a rerun. The second was at ec9b01cb0 (0.41s), 106 packages ok, one FAIL, load 4.9 at the end of the run.
It was taken for a flake, and it is not one. It fails 100% of the time when the process has a single processor, and has never been seen with two or more.
Measured 2026-09-20 on two heads, 1a692da6e (the base) and ec9b01cb0 (the base plus one change to internal/session), with prebuilt test binaries so nothing came from a build cache, the two arms interleaved run by run so any drift on the box fell on both equally:
| Condition | 1a692da6e |
ec9b01cb0 |
|---|---|---|
| single test, unconstrained | 0/20 | 0/20 |
| whole package, unconstrained | 0/8 | 0/8 |
single test, GOMAXPROCS=1 |
30/30 | 30/30 |
single test, GOMAXPROCS=2 |
0/30 | 0/30 |
single test, GOMAXPROCS=4 |
0/30 | 0/30 |
whole package, GOMAXPROCS=1 |
exit 1, this test only | exit 1, this test only |
-race, whole package, ×2 |
pass, 0 data races | pass, 0 data races |
Every one of the 60 failures is byte-identical to the two gate sightings, between 0.36s and 0.39s. The sibling test in the same file, TestARunningRootsFinishDoesNotCloseTheRunOverAnUnreviewedChild, passes at one processor and at two.
The sub-half-second duration is the part that rules out starvation: the test runs under a 30 second context, so a starved box would show as a timeout near 30 seconds, not as a wrong outcome in 380ms. This is an ordering result.
Replication
Deterministic (no model). One command, no keys, no fixtures, about a second:
GOMAXPROCS=1 go test -count=1 -run '^TestAParkedRootsRunReviewsItsChildBeforeItAnswers$' ./internal/run/
Today that prints the block quoted above and exits 1, every time. The control is the same command with GOMAXPROCS=2, which passes every time.
Field (real models). None. No provider, no key and no model is involved; the test scripts its workers.
Where
internal/run/run.go. Supervisor.Run is the loop that answers the outcome, and OutcomeIncomplete is the constant whose text is ran and did not finish. The scenario is in internal/run/review_order_test.go: the supervisor is built with two slots and the test forces an ordering with childDone and lateTried, with no sleeps, so a root parks on its child while a round its worker had already started still ends its tool.
Searchable strings: OutcomeIncomplete, func (s *Supervisor) Run, releaseStale, passInterval, parkedRootWorker.
What has not been established: which line assumes a second processor, and therefore whether the defect is in the supervisor or in the test's own scripting. The scheduling story below is consistent with everything measured and is not proof of a mechanism.
The fix
A run's answer must not depend on how many processors the process was given. Supervisor.Run should answer done for this scripted order on one processor exactly as it does on two, or, if the test is asserting an order the supervisor never promised, the test should assert the promise it actually makes and say so.
Whichever it is, the outcome is decided by the same constants today and no person-facing wording changes, so no manual page is owed by this fix unless the outcome a run reports changes, which it should not.
Acceptance
- e2e:
GOMAXPROCS=1 go test -count=1 ./internal/run/exits 0, with no--- FAILline forTestAParkedRootsRunReviewsItsChildBeforeItAnswers. This is the command that fails today. - e2e: the control,
GOMAXPROCS=4 go test -count=1 ./internal/run/, still exits 0, so the fix does not trade one parallelism for another. - Unit: the scenario pins its own parallelism rather than inheriting the runner's — the test sets
runtime.GOMAXPROCS(1)for its own duration and restores it — so the guarantee is that the run answersdoneon one processor, asserted in the test itself and not dependent on how the gate happens to be configured. Without this the fix is provable only by remembering to pass an environment variable, which is a regression test that can silently stop checking. - And the fix does not read the processor count. Nothing in
internal/runmay branch onruntime.GOMAXPROCSorruntime.NumCPUto satisfy the line above. That branch is the cheapest way to make a pinned scenario pass while leaving the ordering assumption in place for every machine that behaves like one processor under load — which, on the reading above, is how this reached two gates. The shape that counts: the pinned scenario passes, the unconstrained runs keep passing, and the diff contains no read of the processor count.
Two open questions, neither answered here
What parallelism does the gate run under? If a full check runs unconstrained on a many-core box, then two sightings in roughly a dozen full runs is not explained by this result, and the story that load transiently narrows effective parallelism is doing a lot of work. If the gate constrains GOMAXPROCS, the explanation is straightforward. This was not checked.
A lock reading on the test box does not mean the box is quiet. While these runs were taken, two suites were running from other rigs (go test ./internal/tui3 and go test ./internal/session) and the suite flock read FREE throughout, because runs inside cell sandboxes do not take it. Load stayed between 3.2 and 5.6 for the whole experiment. That does not affect this result — 60/60 against 0/120 is not a load effect — but a reader should not take "the lock was free" as "nothing else was running" on this box.
- 主要言語
- Go
- スター
- 115
- フォーク
- 14
- 平均マージ
- 9時間 37分
- マージ済み PR(30日)
- 755
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
Agent-Field/CodeAF のほかの issue
-
area:chat bug sev:papercut
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
Agent-Field/CodeAF#1592 ·
メンテナーはふだん 1 日以内に返信
-
area:headless bug sev:critical
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
Agent-Field/CodeAF#1566 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
area:chat feature
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
Agent-Field/CodeAF#1510 ·
メンテナーはふだん 1 日以内に返信
-
area:tests bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
Agent-Field/CodeAF#1489 ·
メンテナーはふだん 1 日以内に返信
-
area:chat bug good first issue sev:papercut
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
Agent-Field/CodeAF#1470 ·
メンテナーはふだん 1 日以内に返信
Agent-Field/CodeAF の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
rossoctl/context-guru#346 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
prime-radiant-inc/evener#2883 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
gravitational/teleport#69805 ·
メンテナーはふだん 11 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
Under Poisson sampling, the `PLDAccountant` composes the inner event both before and after samplingオープン
難易度 2/5 半日 初心者へのやさしさ 78/100
google/differential-privacy#496 ·