The bashloop benchmark cannot measure its bash arm because it always passes --db
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
Start in bench/bashloop/door.go:169 and runDoDoor, then compare the bash and older-engine invocations with cmd/codeaf/do.go:3430 and runRoadRefusesStore. Run the deterministic reproduction and the argument-table test; done means the bash arm omits -db, the older-engine arm retains it, and the pair benchmark records calls for both without a --db refusal.
由索引模型根据 Issue 内容生成。
描述
Found on santos/dev2 at 008363c98 (#1410). It reaches dev when #1410 merges.
What happened
On 2026-09-24, go run ./bench/bashloop -mode pair -door do completed its older-engine arm but made zero calls in its bash arm. That arm was recorded as could-not-run after codeaf do answered --db names a store only the older engine works in; … Drop --db, or set CODEAF_TASK_BELT=node to run this on the older engine. The benchmark therefore cannot compare the default belt.
Replication
Deterministic (no model). The refusal comes before any model call, so any key value will do:
export CODEAF_HOME=$(mktemp -d) HOME=$(mktemp -d) OPENROUTER_API_KEY=<any non-empty value>
R=$(mktemp -d) && git -C "$R" init -q && git -C "$R" commit -q --allow-empty -m init
env -u CODEAF_TASK_BELT codeaf do -w "$R" -db "$CODEAF_HOME/graph.db" -timeout 60s -yes-spend -json noop; echo "exit $?"
Today it prints the --db names a store only the older engine works in refusal and exit 1. That is the call the benchmark's bash arm makes. With CODEAF_TASK_BELT=node the same command is accepted.
Field (real models). go run ./bench/bashloop -mode pair -door do with OPENROUTER_API_KEY and deepseek/deepseek-v4-flash: a few minutes and a few cents. The older-engine arm makes its calls and passes; the bash arm makes 0 calls and ends could-not-run.
Where
bench/bashloop/door.go:169, runDoDoor, appends -db; cmd/codeaf/do.go:3430, runRoadRefusesStore, rejects it on the bash road.
The fix
Pass -db only to the older-engine arm. Give the bash arm its supported store setup and keep the two arms' receipts comparable.
Acceptance
- e2e:
go run ./bench/bashloop -mode pair -door dorecords at least one call for each arm, with neither markedcould-not-runfor--db. - Unit: an argument-table test asserts the bash invocation omits
-dband the older-engine invocation retains it. - Update benchmark instructions and a change entry's
invalidates.
- 主要语言
- Go
- 星标
- 115
- 派生
- 14
- 平均合并
- 9 小时 44 分钟
- 30 天内合并 PR
- 766
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
Agent-Field/CodeAF 的其他 Issue
-
area:chat bug sev:papercut
难度 2/5 1-3 小时 新手友好度 86/100
Agent-Field/CodeAF#1592 ·
维护者通常 1 天内回复
-
area:headless bug sev:critical
难度 2/5 1-3 小时 新手友好度 88/100
Agent-Field/CodeAF#1566 · 已指派 1 人 ·
维护者通常 1 天内回复
-
area:chat feature
难度 2/5 1-3 小时 新手友好度 88/100
Agent-Field/CodeAF#1510 ·
维护者通常 1 天内回复
-
area:tests bug
难度 2/5 1-3 小时 新手友好度 82/100
Agent-Field/CodeAF#1489 ·
维护者通常 1 天内回复
-
area:chat bug good first issue sev:papercut
难度 2/5 1-3 小时 新手友好度 78/100
Agent-Field/CodeAF#1470 ·
维护者通常 1 天内回复
查看 Agent-Field/CodeAF 的全部 Issue
相似的 Issue
-
agentic-workflows
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
维护者通常 1 天内回复
-
priority/4/normal status/needs-triage type/bug/unconfirmed
难度 2/5 1-3 小时 新手友好度 78/100
authelia/authelia#13292 · 1 条评论 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
blinklabs-io/actions#138 ·
维护者通常 1 天内回复
-
[UI] AlbumDetails collapses multi-genre list to single primary genre on viewports < lg breakpoint未关闭
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复