Flaky e2e: gateway-fleet reads a worker's capacity before the worker has left starting
维护者通常 1 天内回复
评估
- 难度
- 4/5
- 预计耗时
- 1-2 天
- 新手友好度
- 48/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
- 技术栈
- typescript
调研方向
Start in e2e/gateway-fleet.test.ts at waitForWorkers and iosLimit, then read src/gateway/worker-link.ts (#rebuildView, #refreshStarting) and the changesCapacityOrLeases filter to see why daemon.started never triggers a refresh. Triage must settle the product question (refresh on leaving starting) before the test fix, and an open PR (#421) already claims the work; done means waitForWorkers waits for capacity/catalog presence and CI passes under the CPU-load conditions described.
由索引模型根据 Issue 内容生成。
描述
What fails
e2e/gateway-fleet.test.ts > "gateway fleet > two workers join, status aggregates them, a kill disconnects one, drain flags the other" fails with AssertionError: expected 0 to be greater than 0 at expect(iosLimit(joined[0])).toBeGreaterThan(0). Seen once in CI on PR #409 (run 37477061691).
Cause: a race on main, not in a PR
The test waits only for both workers to be connected, then reads capacity from that first view. A worker that dials the gateway while its daemon is still starting answers status.get with health and host only (#398), so the gateway builds its view without capacity (worker-link.ts #rebuildView -> #refreshStarting). The view stays without capacity until the next refresh, and none is triggered when the worker becomes running: daemon.started is not a lease./device./install event (changesCapacityOrLeases), so the next read is a device or lease event or the 30s periodic tick. The test's iosLimit falls back to 0.
Evidence
- On
origin/main(37188b5), unmodified: 6 of 6 runs of the file pass under CPU load (16 busy loops on 12 cores), 5 of 5 unloaded. - On
origin/mainwithawait new Promise((r) => setTimeout(r, 1500))as the first line ofconvergeStartupinsrc/daemon/main.ts(a slower startup): 3 of 3 runs fail with this exact assertion (2 tests failed each time). - The window is the worker's startup time, so anything that lengthens startup (PR #409 adds a serial read, doctor pass and lease reconcile) makes it likelier. It is not specific to that PR.
Two things to decide in triage
- Product: should the gateway refresh a worker's view as soon as it leaves
starting? Today a freshly joined worker takes no routed work for up to 30s after its startup ends if no device or lease event happens first (checkFleetCoordinator/routing on a view with no catalog). - Test:
waitForWorkersin the file should wait forcapacity(and catalog) to be present before the test reads it.
- 主要语言
- TypeScript
- 星标
- 15
- 派生
- 1
- 平均合并
- 8 小时 23 分钟
- 30 天内合并 PR
- 136
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
callstackincubator/simlock 的其他 Issue
-
bug:new
难度 2/5 1-3 小时 新手友好度 78/100
callstackincubator/simlock#424 ·
维护者通常 1 天内回复
-
bug:new
难度 2/5 1-3 小时 新手友好度 84/100
callstackincubator/simlock#422 ·
维护者通常 1 天内回复
-
bug:new
难度 1/5 1 小时以内 新手友好度 88/100
callstackincubator/simlock#420 ·
维护者通常 1 天内回复
-
bug:new
难度 2/5 1-3 小时 新手友好度 76/100
callstackincubator/simlock#350 · 1 条评论 ·
维护者通常 1 天内回复
-
feature:spec
难度 4/5 3-5 天 新手友好度 35/100
callstackincubator/simlock#412 ·
维护者通常 1 天内回复
查看 callstackincubator/simlock 的全部 Issue
相似的 Issue
-
ble-needs-fable-review bug mobile priority:P2
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 76/100
ColeMurray/background-agents#2305 ·
维护者通常 1 天内回复
-
bug from-studio
难度 2/5 1-3 小时 新手友好度 63/100
esengine/DeepSeek-Reasonix#12355 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
oblien/openship#1086 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复