`eval_clean_hashes` is missing or stale for some `task_cleanroom_v6` images
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 58/100
Research direction
Start with src/programbench/data/tasks/*/task.yaml, eval_batch.py, and Evaluator._remove_hashed_files(). Run the provided docker sha256sum command for the affected task_cleanroom_v6 images and compare each digest with its metadata. Done means every real task has current, well-formed deduplicated hashes and the update process is documented or automated for rebuilt images.
Written by the indexing model from the issue text.
Description
Summary
eval_clean_hashes is intended to remove byte-for-byte copies of the gold executable from a submission before compile.sh runs. However, some task metadata does not define the field, and some existing values do not match the reference executable currently shipped in the corresponding task_cleanroom_v6 image.
This makes the metadata incomplete for downstream evaluators that rely on eval_clean_hashes to scrub renamed copies of the reference executable.
Current state
At repository commit 963063c:
- 201 task directories total (including
testorg__calculator.abc1234) - 40 task YAML files do not define
eval_clean_hashes - 161 define at least one hash
- none define an explicitly empty list
Excluding the test fixture, 39 of the 200 real tasks are missing the field.
In addition, a non-empty list does not necessarily contain the SHA-256 of the executable in the current task_cleanroom_v6 image.
Examples observed:
| Task | Hash in task.yaml |
Current task_cleanroom_v6 reference SHA-256 |
|---|---|---|
facebookresearch__fasttext.1142dc4 |
missing | 80310c0ae4d92165b7a94b17439a1931d99567ae4c1d54706bb557e87bdcd51c |
halitechallenge__halite.822cfb6 |
missing | 9d5e3ec763bbfa7c1091054c3a746ee25af7d4368edc4395eb3e3d8f7f9feac3 |
tomnomnom__gron.88a6234 |
missing | 0270943958c042821654a1e19608848a8cf14cd1136fde4726e8a1821cdf4cc7 |
rs__curlie.5dfcbb1 |
e7e4f84ba194306ef782853d8b77788dd4674a6baad19c878885bb08c2db0d35 |
7ebf850e7d8d0def143b72a1d6d0ebf4e54d8212c515c2317f06e4e8f3f642f0 |
wfxr__csview.8ac4de0 |
e07420f640d1e6a0c047cab49431d1b561f661241efe111a053ff5dde1119d09 |
758ef03ba091d6a7c77d32e60aaf484aca23532aba3eea6a51bdf25232c5f5ce |
As a control, xorg62__tty-clock.f2f847c does contain the current reference hash (cd400708...).
Reproduction
For a task image:
docker run --rm \
--network none \
--user root \
--entrypoint sha256sum \
programbench/rs_1776_curlie.5dfcbb1:task_cleanroom_v6 \
/workspace/executable
Compare the resulting digest with:
src/programbench/data/tasks/rs__curlie.5dfcbb1/task.yaml
The evaluator consumes this metadata in eval_batch.py and recursively removes matching files in Evaluator._remove_hashed_files() before compilation. The evaluator comment describes these hashes as protection against byte-for-byte copies of the gold binary, so including the current cleanroom reference hash appears consistent with the existing field semantics.
Suggested fix
- Compute the SHA-256 of
/workspace/executablefor every currenttask_cleanroom_v6image. - Append it to that task's
eval_clean_hashes, preserving historical hashes and deduplicating the list. - Add a maintenance command or release script that updates these values whenever cleanroom images are rebuilt.
- Validate that every real task has at least one well-formed 64-character SHA-256 value.
- Ideally pin cleanroom images by digest, or verify/update the hashes as part of the image publication workflow, so mutable tags cannot silently drift from task metadata.
I can prepare a follow-up PR to populate the current hashes if this approach matches the intended use of eval_clean_hashes.
- Dominant language
- Python
- Stars
- 928
- Forks
- 67
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from facebookresearch/ProgramBench
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
-
Difficulty 4/5 3-5 days Newbie friendliness 55/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
facebookresearch/ProgramBench#50 · 1 comment ·
All issues in facebookresearch/ProgramBench
Similar issues
-
area: harness bug status: needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Human-Agent-Society/reef#625 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
Difficulty 1/5 Under an hour Newbie friendliness 80/100
learningequality/kolibri#15351 · 2 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Name consistency Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
eellak/triplestore#65 · 1 comment ·