huggingface/lighteval

[EVAL] Big-Bench Extra Hard (BBEH)

Aperta

#600 aperta il 3 mar 2025

 (3 commenti) (0 reazioni) (0 assegnatari)Python (514 fork)auto 404
good first issuehelp wantednew-taskscience-team

Metriche repository

Star
 (2496 stelle)
Metriche merge PR
 (Metriche PR in attesa)

Descrizione

Evaluation short description

Google has releases BBEH as a way to compensate for the saturation of BBH in the latest generation of LLMs. Overall looks like a good benchmark to probe reasoning capabilities.

Evaluation metadata

Provide all available

Guida contributor