`eval_clean_hashes` is missing or stale for some `task_cleanroom_v6` images
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 58/100
調査の方向性
src/programbench/data/tasks/*/task.yaml、eval_batch.py、Evaluator._remove_hashed_files()から始めます。影響を受けるtask_cleanroom_v6イメージに対して、提供されているdocker sha256sumコマンドを実行し、各ダイジェストをメタデータと比較します。すべての実タスクに最新かつ適切な形式の重複排除済みハッシュがあり、再ビルドされたイメージの更新プロセスが文書化または自動化されていれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Summary
eval_clean_hashes is intended to remove byte-for-byte copies of the gold executable from a submission before compile.sh runs. However, some task metadata does not define the field, and some existing values do not match the reference executable currently shipped in the corresponding task_cleanroom_v6 image.
This makes the metadata incomplete for downstream evaluators that rely on eval_clean_hashes to scrub renamed copies of the reference executable.
Current state
At repository commit 963063c:
- 201 task directories total (including
testorg__calculator.abc1234) - 40 task YAML files do not define
eval_clean_hashes - 161 define at least one hash
- none define an explicitly empty list
Excluding the test fixture, 39 of the 200 real tasks are missing the field.
In addition, a non-empty list does not necessarily contain the SHA-256 of the executable in the current task_cleanroom_v6 image.
Examples observed:
| Task | Hash in task.yaml |
Current task_cleanroom_v6 reference SHA-256 |
|---|---|---|
facebookresearch__fasttext.1142dc4 |
missing | 80310c0ae4d92165b7a94b17439a1931d99567ae4c1d54706bb557e87bdcd51c |
halitechallenge__halite.822cfb6 |
missing | 9d5e3ec763bbfa7c1091054c3a746ee25af7d4368edc4395eb3e3d8f7f9feac3 |
tomnomnom__gron.88a6234 |
missing | 0270943958c042821654a1e19608848a8cf14cd1136fde4726e8a1821cdf4cc7 |
rs__curlie.5dfcbb1 |
e7e4f84ba194306ef782853d8b77788dd4674a6baad19c878885bb08c2db0d35 |
7ebf850e7d8d0def143b72a1d6d0ebf4e54d8212c515c2317f06e4e8f3f642f0 |
wfxr__csview.8ac4de0 |
e07420f640d1e6a0c047cab49431d1b561f661241efe111a053ff5dde1119d09 |
758ef03ba091d6a7c77d32e60aaf484aca23532aba3eea6a51bdf25232c5f5ce |
As a control, xorg62__tty-clock.f2f847c does contain the current reference hash (cd400708...).
Reproduction
For a task image:
docker run --rm \
--network none \
--user root \
--entrypoint sha256sum \
programbench/rs_1776_curlie.5dfcbb1:task_cleanroom_v6 \
/workspace/executable
Compare the resulting digest with:
src/programbench/data/tasks/rs__curlie.5dfcbb1/task.yaml
The evaluator consumes this metadata in eval_batch.py and recursively removes matching files in Evaluator._remove_hashed_files() before compilation. The evaluator comment describes these hashes as protection against byte-for-byte copies of the gold binary, so including the current cleanroom reference hash appears consistent with the existing field semantics.
Suggested fix
- Compute the SHA-256 of
/workspace/executablefor every currenttask_cleanroom_v6image. - Append it to that task's
eval_clean_hashes, preserving historical hashes and deduplicating the list. - Add a maintenance command or release script that updates these values whenever cleanroom images are rebuilt.
- Validate that every real task has at least one well-formed 64-character SHA-256 value.
- Ideally pin cleanroom images by digest, or verify/update the hashes as part of the image publication workflow, so mutable tags cannot silently drift from task metadata.
I can prepare a follow-up PR to populate the current hashes if this approach matches the intended use of eval_clean_hashes.
- 主要言語
- Python
- スター
- 928
- フォーク
- 67
- PR マージ指標
- 30日以内にマージされた PR はありません
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
facebookresearch/ProgramBench のほかの issue
-
難易度 3/5 1〜2日 初心者へのやさしさ 68/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
facebookresearch/ProgramBench#50 · コメント 1 件 ·
facebookresearch/ProgramBench の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
anthropics/skills#1811 · コメント 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
speaches-ai/speaches#678 ·
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
datalayer/mcp-compose#42 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
conda-forge/spacy-feedstock#177 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
UKGovernmentBEIS/inspect_evals#2523 ·