Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

ProgramBench csview: false not_run records and missing results.xml

Open
#56 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Quiet
Tech stack
docker, python
Domain
testing, tooling

Research direction

Start by inspecting the attached raw evaluation JSON, audit script, and technical report, then compare tests.json with the JUnit names and inspect the missing results.xml case. Check the help.txt reference against the preserved executable. Done means matching tests are counted once, failures expose runner output and emit usable results, and the help golden is tied to the stated reference.

Written by the indexing model from the issue text.

Description

Environment
  • ProgramBench: 1.2.4
  • Repository commit: 963063c9271cc40fa179977356782ea4582e0b0c
  • Instance: wfxr__csview.8ac4de0
  • Cleanroom image: programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6
  • Image digest: sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd
  • Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
  • Inference container: --network none
Reproduction

The same submission was evaluated twice:

uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5

uv run programbench eval \
  ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
  --docker-cpus 1 \
  --branch-retries 3

Both runs produced the same raw result counts:

  • passed: 344
  • failed: 1
  • skipped: 1
  • not_run: 155
  • displayed score: 68

The generated executable SHA-256 was identical in both evaluations:

a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888

Issue 1: test-name prefix mismatch creates 153 false not_run records

For branch d31025219260, the evaluator reports both:

  • 153/153 expected tests missing from JUnit XML
  • 153 test(s) in JUnit XML not in tests.json

The records form 153 exact pairs after removing only the leading eval. prefix. For example:

  • metadata: tests.test_edge_cases.test_all_cells_empty
  • JUnit: eval.tests.test_edge_cases.test_all_cells_empty

All 153 JUnit-side records passed, while the metadata-side names were counted as not_run.

Suggested fix: normalize an optional leading eval. prefix before matching JUnit names to tests.json, or generate both sources with the same package root.

Issue 2: branch efa8c407dbe3 never produces results.xml

The branch fails with:

results_read_failed: Could not find the file /workspace/eval/results.xml

This occurred in the default run and again in the serial run. With --branch-retries 3, the initial attempt plus all three retries failed identically. Two expected tests remain not_run because no JUnit result file is available.

Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when results.xml is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.

Issue 3: --help golden differs from the cleanroom reference behavior

The only completed-test failure is:

eval.tests.test_cli_basics.test_help_exact

However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:

diff -u \
  <(/workspace/executable --help) \
  <(/candidate/csview_reconstruction_v0_5/executable --help)

The local cleanroom differential suite also passed 49/49, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in task_cleanroom_v6.

Suggested fix: regenerate help.txt directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.

Why this matters

Among completed scored tests, the observed pass rate is 344 / 345 = 99.710145%. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.

This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.

Public evidence attached

The attached public-safe evidence archive contains:

  • both raw evaluation JSON files;
  • a machine-readable audit summary;
  • the deterministic audit script;
  • the full bilingual technical report;
  • SHA-256 checksums.

No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.

programbench_csview_public_issue_evidence_v1.zip

Dominant language
Python
Stars
928
Forks
67
PR merge metrics
No merged PRs in 30d

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from facebookresearch/ProgramBench

All issues in facebookresearch/ProgramBench

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.