huggingface/lighteval

[EVAL] Add kyrgyzLLM benchmark

開放

#1,036 建立於 2025年11月4日

 (1 則留言) (0 個反應) (1 位負責人)Python (514 個分叉)auto 404
good first issuenew-task

倉庫指標

星標
 (2,496 顆星)
PR 合併指標
 (PR 指標待抓取)

描述

Hi,

We just open-sourced the Kyrgyz LLM Evaluation Dataset.

Evaluation short description

  • Why is this evaluation interesting?

KyrgyzLLM-Bench is the first comprehensive benchmark suite for deep language understanding in Kyrgyz. It is interesting because it provides broad, culturally grounded coverage by combining native benchmarks (such as KyrgyzMMLU and KyrgyzRC) with carefully translated and post-edited international benchmarks (such as HellaSwag, WinoGrande, BoolQ, GSM8K, and TruthfulQA).

  • How used is it in the community? As the benchmark was released recently, its adoption by the community is just beginning. It is significant because it's the first comprehensive benchmark suite for deep language understanding, specifically in the Kyrgyz language, providing a new and essential tool for researchers and developers.

Evaluation metadata

Thanks you!

貢獻者指南