huggingface/lighteval

[EVAL] Big-Bench Extra Hard (BBEH)

オープン

#600 opened on 2025/03/03

 (3 件のコメント) (0 件のリアクション) (0 人の担当者)Python (514 件のフォーク)auto 404
good first issuehelp wantednew-taskscience-team

Repository metrics

Stars
 (2,496 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Evaluation short description

Google has releases BBEH as a way to compensate for the saturation of BBH in the latest generation of LLMs. Overall looks like a good benchmark to probe reasoning capabilities.

Evaluation metadata

Provide all available

コントリビューターガイド