huggingface/lighteval

[EVAL] Big-Bench Extra Hard (BBEH)

开放

#600 创建于 2025年3月3日

 (3 条评论) (0 个反应) (0 位负责人)Python (514 个派生)auto 404
good first issuehelp wantednew-taskscience-team

仓库指标

星标
 (2,496 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

Evaluation short description

Google has releases BBEH as a way to compensate for the saturation of BBH in the latest generation of LLMs. Overall looks like a good benchmark to probe reasoning capabilities.

Evaluation metadata

Provide all available

贡献者指南