Request gold compile.sh scripts
还没有人认领这个 Issue。
评估
调研方向
找到 gold/reference compile.sh 脚本,并检查评估流水线如何使用 repository 提交和已构建的可执行文件。确定是否可以在不损害 benchmark 完整性的情况下发布这些脚本;完成的标准是提供 reference build 脚本,或记录一种受支持的创建已知有效 mock submissions 的方法。
由索引模型根据 Issue 内容生成。
描述
Hi there.
Really like this benchmark. I’m building a Modal-based pipeline to evaluate ProgramBench with higher parallelism, and I’d like to validate my implementation end-to-end without running a full agent each time.
For that purpose, I’m looking for a way to mock the agent generation phase with known-good submissions. The repo part is simple as the commit hash is known. As for the compile script part, it is a little bit tricky. Would it be possible to release the compile.sh scripts used to build the gold/reference executables, or any equivalent reference build scripts?
I understand if these cannot be shared due to benchmark integrity concerns. In that case, is there a recommended way to create a small set of known-good mock submissions for validating the evaluation pipeline?
- 主要语言
- Python
- 星标
- 928
- 派生
- 67
- PR 合并指标
- 30 天内没有已合并 PR
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
facebookresearch/ProgramBench 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 68/100
-
难度 4/5 3-5 天 新手友好度 55/100
-
难度 4/5 3-5 天 新手友好度 35/100
-
难度 4/5 3-5 天 新手友好度 48/100
-
难度 4/5 3-5 天 新手友好度 58/100
查看 facebookresearch/ProgramBench 的全部 Issue
相似的 Issue
-
area: harness bug status: needs-triage
难度 2/5 1-3 小时 新手友好度 75/100
Human-Agent-Society/reef#625 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 1/5 1 小时以内 新手友好度 80/100
learningequality/kolibri#15351 · 2 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
-
Name consistency 未关闭
难度 2/5 1-3 小时 新手友好度 75/100
eellak/triplestore#65 · 1 条评论 ·