Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

ProgramBench csview: false not_run records and missing results.xml

未关闭
#56 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
48/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
docker, python
领域
testing, tooling

调研方向

首先检查随附的原始评估 JSON、审计脚本和技术报告,然后将 tests.json 与 JUnit 名称进行比较,并检查缺少 results.xml 的情况。将 help.txt 中的参考与保留的可执行文件进行核对。当匹配的测试只被计数一次、失败会暴露 runner 输出并生成可用结果,且 help golden 与所述参考绑定时,即视为完成。

由索引模型根据 Issue 内容生成。

描述

Environment
  • ProgramBench: 1.2.4
  • Repository commit: 963063c9271cc40fa179977356782ea4582e0b0c
  • Instance: wfxr__csview.8ac4de0
  • Cleanroom image: programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6
  • Image digest: sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd
  • Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
  • Inference container: --network none
Reproduction

The same submission was evaluated twice:

uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5

uv run programbench eval \
  ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
  --docker-cpus 1 \
  --branch-retries 3

Both runs produced the same raw result counts:

  • passed: 344
  • failed: 1
  • skipped: 1
  • not_run: 155
  • displayed score: 68

The generated executable SHA-256 was identical in both evaluations:

a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888

Issue 1: test-name prefix mismatch creates 153 false not_run records

For branch d31025219260, the evaluator reports both:

  • 153/153 expected tests missing from JUnit XML
  • 153 test(s) in JUnit XML not in tests.json

The records form 153 exact pairs after removing only the leading eval. prefix. For example:

  • metadata: tests.test_edge_cases.test_all_cells_empty
  • JUnit: eval.tests.test_edge_cases.test_all_cells_empty

All 153 JUnit-side records passed, while the metadata-side names were counted as not_run.

Suggested fix: normalize an optional leading eval. prefix before matching JUnit names to tests.json, or generate both sources with the same package root.

Issue 2: branch efa8c407dbe3 never produces results.xml

The branch fails with:

results_read_failed: Could not find the file /workspace/eval/results.xml

This occurred in the default run and again in the serial run. With --branch-retries 3, the initial attempt plus all three retries failed identically. Two expected tests remain not_run because no JUnit result file is available.

Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when results.xml is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.

Issue 3: --help golden differs from the cleanroom reference behavior

The only completed-test failure is:

eval.tests.test_cli_basics.test_help_exact

However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:

diff -u \
  <(/workspace/executable --help) \
  <(/candidate/csview_reconstruction_v0_5/executable --help)

The local cleanroom differential suite also passed 49/49, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in task_cleanroom_v6.

Suggested fix: regenerate help.txt directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.

Why this matters

Among completed scored tests, the observed pass rate is 344 / 345 = 99.710145%. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.

This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.

Public evidence attached

The attached public-safe evidence archive contains:

  • both raw evaluation JSON files;
  • a machine-readable audit summary;
  • the deterministic audit script;
  • the full bilingual technical report;
  • SHA-256 checksums.

No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.

programbench_csview_public_issue_evidence_v1.zip

主要语言
Python
星标
928
派生
67
PR 合并指标
30 天内没有已合并 PR

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

facebookresearch/ProgramBench 的其他 Issue

查看 facebookresearch/ProgramBench 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。