ProgramBench csview: false not_run records and missing results.xml
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
Direzione di ricerca
Inizia ispezionando il JSON di valutazione grezzo allegato, lo script di audit e il rapporto tecnico, quindi confronta tests.json con i nomi JUnit e esamina il caso di results.xml mancante. Verifica il riferimento in help.txt rispetto all’eseguibile conservato. Il lavoro è completato quando i test corrispondenti vengono conteggiati una sola volta, i fallimenti espongono l’output del runner e producono risultati utilizzabili, e il golden di help è associato al riferimento dichiarato.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Environment
- ProgramBench:
1.2.4 - Repository commit:
963063c9271cc40fa179977356782ea4582e0b0c - Instance:
wfxr__csview.8ac4de0 - Cleanroom image:
programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6 - Image digest:
sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd - Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
- Inference container:
--network none
Reproduction
The same submission was evaluated twice:
uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5
uv run programbench eval \
~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
--docker-cpus 1 \
--branch-retries 3
Both runs produced the same raw result counts:
- passed: 344
- failed: 1
- skipped: 1
- not_run: 155
- displayed score: 68
The generated executable SHA-256 was identical in both evaluations:
a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888
Issue 1: test-name prefix mismatch creates 153 false not_run records
For branch d31025219260, the evaluator reports both:
153/153 expected tests missing from JUnit XML153 test(s) in JUnit XML not in tests.json
The records form 153 exact pairs after removing only the leading eval. prefix. For example:
- metadata:
tests.test_edge_cases.test_all_cells_empty - JUnit:
eval.tests.test_edge_cases.test_all_cells_empty
All 153 JUnit-side records passed, while the metadata-side names were counted as not_run.
Suggested fix: normalize an optional leading eval. prefix before matching JUnit names to tests.json, or generate both sources with the same package root.
Issue 2: branch efa8c407dbe3 never produces results.xml
The branch fails with:
results_read_failed: Could not find the file /workspace/eval/results.xml
This occurred in the default run and again in the serial run. With --branch-retries 3, the initial attempt plus all three retries failed identically. Two expected tests remain not_run because no JUnit result file is available.
Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when results.xml is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.
Issue 3: --help golden differs from the cleanroom reference behavior
The only completed-test failure is:
eval.tests.test_cli_basics.test_help_exact
However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:
diff -u \
<(/workspace/executable --help) \
<(/candidate/csview_reconstruction_v0_5/executable --help)
The local cleanroom differential suite also passed 49/49, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in task_cleanroom_v6.
Suggested fix: regenerate help.txt directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.
Why this matters
Among completed scored tests, the observed pass rate is 344 / 345 = 99.710145%. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.
This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.
Public evidence attached
The attached public-safe evidence archive contains:
- both raw evaluation JSON files;
- a machine-readable audit summary;
- the deterministic audit script;
- the full bilingual technical report;
- SHA-256 checksums.
No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.
- Lingua principale
- Python
- Stelle
- 928
- Fork
- 67
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di facebookresearch/ProgramBench
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 68/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 58/100
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
facebookresearch/ProgramBench#50 · 1 commento ·
Tutte le issue di facebookresearch/ProgramBench
Issue simili
-
bug confirmed issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
open-webui/open-webui#30750 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
OpenwaterHealth/openmotion-bloodflow-app#604 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
good first issue
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100