huggingface/lighteval

[EVAL] Big-Bench Extra Hard (BBEH)

開放

#600 建立於 2025年3月3日

 (3 則留言) (0 個反應) (0 位負責人)Python (514 個分叉)auto 404
good first issuehelp wantednew-taskscience-team

倉庫指標

星標
 (2,496 顆星)
PR 合併指標
 (PR 指標待抓取)

描述

Evaluation short description

Google has releases BBEH as a way to compensate for the saturation of BBH in the latest generation of LLMs. Overall looks like a good benchmark to probe reasoning capabilities.

Evaluation metadata

Provide all available

貢獻者指南