ProgramBench csview: false not_run records and missing results.xml
还没有人认领这个 Issue。
评估
调研方向
首先检查随附的原始评估 JSON、审计脚本和技术报告,然后将 tests.json 与 JUnit 名称进行比较,并检查缺少 results.xml 的情况。将 help.txt 中的参考与保留的可执行文件进行核对。当匹配的测试只被计数一次、失败会暴露 runner 输出并生成可用结果,且 help golden 与所述参考绑定时,即视为完成。
由索引模型根据 Issue 内容生成。
描述
Environment
- ProgramBench:
1.2.4 - Repository commit:
963063c9271cc40fa179977356782ea4582e0b0c - Instance:
wfxr__csview.8ac4de0 - Cleanroom image:
programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6 - Image digest:
sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd - Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
- Inference container:
--network none
Reproduction
The same submission was evaluated twice:
uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5
uv run programbench eval \
~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
--docker-cpus 1 \
--branch-retries 3
Both runs produced the same raw result counts:
- passed: 344
- failed: 1
- skipped: 1
- not_run: 155
- displayed score: 68
The generated executable SHA-256 was identical in both evaluations:
a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888
Issue 1: test-name prefix mismatch creates 153 false not_run records
For branch d31025219260, the evaluator reports both:
153/153 expected tests missing from JUnit XML153 test(s) in JUnit XML not in tests.json
The records form 153 exact pairs after removing only the leading eval. prefix. For example:
- metadata:
tests.test_edge_cases.test_all_cells_empty - JUnit:
eval.tests.test_edge_cases.test_all_cells_empty
All 153 JUnit-side records passed, while the metadata-side names were counted as not_run.
Suggested fix: normalize an optional leading eval. prefix before matching JUnit names to tests.json, or generate both sources with the same package root.
Issue 2: branch efa8c407dbe3 never produces results.xml
The branch fails with:
results_read_failed: Could not find the file /workspace/eval/results.xml
This occurred in the default run and again in the serial run. With --branch-retries 3, the initial attempt plus all three retries failed identically. Two expected tests remain not_run because no JUnit result file is available.
Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when results.xml is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.
Issue 3: --help golden differs from the cleanroom reference behavior
The only completed-test failure is:
eval.tests.test_cli_basics.test_help_exact
However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:
diff -u \
<(/workspace/executable --help) \
<(/candidate/csview_reconstruction_v0_5/executable --help)
The local cleanroom differential suite also passed 49/49, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in task_cleanroom_v6.
Suggested fix: regenerate help.txt directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.
Why this matters
Among completed scored tests, the observed pass rate is 344 / 345 = 99.710145%. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.
This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.
Public evidence attached
The attached public-safe evidence archive contains:
- both raw evaluation JSON files;
- a machine-readable audit summary;
- the deterministic audit script;
- the full bilingual technical report;
- SHA-256 checksums.
No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.
- 主要语言
- Python
- 星标
- 928
- 派生
- 67
- PR 合并指标
- 30 天内没有已合并 PR
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
facebookresearch/ProgramBench 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 68/100
-
难度 4/5 3-5 天 新手友好度 55/100
-
难度 4/5 3-5 天 新手友好度 35/100
-
难度 4/5 3-5 天 新手友好度 58/100
-
难度 5/5 一周以上 新手友好度 25/100
facebookresearch/ProgramBench#50 · 1 条评论 ·
查看 facebookresearch/ProgramBench 的全部 Issue
相似的 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
stephrobert/dsoxlab#238 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
-
难度 2/5 1-3 小时 新手友好度 75/100
sublimehq/package_control#1780 ·
-
难度 2/5 1-3 小时 新手友好度 65/100
-
难度 2/5 1-3 小时 新手友好度 70/100
nwg-piotr/nwg-displays#145 ·