huggingface/lighteval

[EVAL] Big-Bench Extra Hard (BBEH)

Offen

#600 geöffnet am 03.03.2025

 (3 Kommentare) (0 Reaktionen) (0 zugewiesene Personen)Python (514 Forks)auto 404
good first issuehelp wantednew-taskscience-team

Repository-Metriken

Stars
 (2.496 Sterne)
PR-Merge-Metriken
 (PR-Metriken ausstehend)

Beschreibung

Evaluation short description

Google has releases BBEH as a way to compensate for the saturation of BBH in the latest generation of LLMs. Overall looks like a good benchmark to probe reasoning capabilities.

Evaluation metadata

Provide all available

Contributor Guide