Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

ProgramBench csview: false not_run records and missing results.xml

Aperta
#56 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
48/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
docker, python
Ambito
testing, tooling

Direzione di ricerca

Inizia ispezionando il JSON di valutazione grezzo allegato, lo script di audit e il rapporto tecnico, quindi confronta tests.json con i nomi JUnit e esamina il caso di results.xml mancante. Verifica il riferimento in help.txt rispetto all’eseguibile conservato. Il lavoro è completato quando i test corrispondenti vengono conteggiati una sola volta, i fallimenti espongono l’output del runner e producono risultati utilizzabili, e il golden di help è associato al riferimento dichiarato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Environment
  • ProgramBench: 1.2.4
  • Repository commit: 963063c9271cc40fa179977356782ea4582e0b0c
  • Instance: wfxr__csview.8ac4de0
  • Cleanroom image: programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6
  • Image digest: sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd
  • Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
  • Inference container: --network none
Reproduction

The same submission was evaluated twice:

uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5

uv run programbench eval \
  ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
  --docker-cpus 1 \
  --branch-retries 3

Both runs produced the same raw result counts:

  • passed: 344
  • failed: 1
  • skipped: 1
  • not_run: 155
  • displayed score: 68

The generated executable SHA-256 was identical in both evaluations:

a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888

Issue 1: test-name prefix mismatch creates 153 false not_run records

For branch d31025219260, the evaluator reports both:

  • 153/153 expected tests missing from JUnit XML
  • 153 test(s) in JUnit XML not in tests.json

The records form 153 exact pairs after removing only the leading eval. prefix. For example:

  • metadata: tests.test_edge_cases.test_all_cells_empty
  • JUnit: eval.tests.test_edge_cases.test_all_cells_empty

All 153 JUnit-side records passed, while the metadata-side names were counted as not_run.

Suggested fix: normalize an optional leading eval. prefix before matching JUnit names to tests.json, or generate both sources with the same package root.

Issue 2: branch efa8c407dbe3 never produces results.xml

The branch fails with:

results_read_failed: Could not find the file /workspace/eval/results.xml

This occurred in the default run and again in the serial run. With --branch-retries 3, the initial attempt plus all three retries failed identically. Two expected tests remain not_run because no JUnit result file is available.

Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when results.xml is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.

Issue 3: --help golden differs from the cleanroom reference behavior

The only completed-test failure is:

eval.tests.test_cli_basics.test_help_exact

However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:

diff -u \
  <(/workspace/executable --help) \
  <(/candidate/csview_reconstruction_v0_5/executable --help)

The local cleanroom differential suite also passed 49/49, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in task_cleanroom_v6.

Suggested fix: regenerate help.txt directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.

Why this matters

Among completed scored tests, the observed pass rate is 344 / 345 = 99.710145%. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.

This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.

Public evidence attached

The attached public-safe evidence archive contains:

  • both raw evaluation JSON files;
  • a machine-readable audit summary;
  • the deterministic audit script;
  • the full bilingual technical report;
  • SHA-256 checksums.

No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.

programbench_csview_public_issue_evidence_v1.zip

Lingua principale
Python
Stelle
928
Fork
67
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di facebookresearch/ProgramBench

Tutte le issue di facebookresearch/ProgramBench

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.