Issues du dépôt
huggingface/lighteval
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Issues
Ouverte
[EVAL] Long Horizon Execution
good first issuehelp wantednew-task
Label adapté aux débutants
7 commentaires0 réaction1 personne assignée
Ouverte
[EVAL] Add kyrgyzLLM benchmark
good first issuenew-task
Pourquoi recommandéeLabel adapté aux débutants
Label adapté aux débutants
1 commentaire0 réaction1 personne assignée
Ouverte
[FT] showing count in Markdown summary table
featuregood first issue
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
5 commentaires1 réaction0 personne assignée
Ouverte
[FT] Add tests for nanotron
featuregood first issuehelp wantedignore-for-release
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
3 commentaires0 réaction0 personne assignée
Ouverte
[BUG] custom model docs don't run: missing imports
bugdocumentationgood first issue
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
5 commentaires0 réaction0 personne assignée
Ouverte
[FT] Manage script and language in the Language enum
featuregood first issue
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
2 commentaires0 réaction0 personne assignée
Ouverte
[EVAL] TauBench:
help wantednew-task
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
1 commentaire0 réaction0 personne assignée
Ouverte
[EVAL] SciCode: reasearch coding benchmark
help wantednew-task
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
2 commentaires0 réaction0 personne assignée
Ouverte
[BUG] Optimize tokenization
buggood first issuehelp wantedscience-team
Pourquoi recommandéeLabel adapté aux débutants
Label adapté aux débutants
1 commentaire0 réaction1 personne assignée
Ouverte
[EVAL] HELMET: long context evals
help wantednew-task
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
2 commentaires0 réaction0 personne assignée
Ouverte
[EVAL] SWEBENCH multilingual
help wantednew-task
Pourquoi recommandéeAucune personne assignée · Aucun commentaire pour l'instant
Aucune personne assignéeAucun commentaire pour l'instantLabel adapté aux débutants
0 commentaire0 réaction0 personne assignée
Ouverte
[FT] Add tests for `VLLMModel` base methods
featuregood first issue
Pourquoi recommandéeAucun commentaire pour l'instant · Label adapté aux débutants
Aucun commentaire pour l'instantLabel adapté aux débutants
0 commentaire1 réaction1 personne assignée
Ouverte
Call for contributions: Translate lighteval's doc into Chinese
good first issuehelp wantednew-task
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
2 commentaires1 réaction0 personne assignée
Ouverte
[BUG] Installing lighteval breaks hydra-core
bughelp wanted
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
1 commentaire1 réaction0 personne assignée
Ouverte
[EVAL] Adding PHARE
good first issuehelp wantednew-task
Pourquoi recommandéeLabel adapté aux débutants
Label adapté aux débutants
9 commentaires0 réaction1 personne assignée
Ouverte
[FT] Improve Documentation and Examples
documentationfeaturegood first issuehelp wanted
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
6 commentaires1 réaction0 personne assignée
Ouverte
[FT] Log progress bar on main process
featurehelp wantedscience-team
Pourquoi recommandéeAucune personne assignée · Aucun commentaire pour l'instant
Aucune personne assignéeAucun commentaire pour l'instantLabel adapté aux débutants
0 commentaire0 réaction0 personne assignée
Ouverte
[FT] Build in a way to specify specific IDs/Lines in Dataset to use as few-shot examples in the same split
featuregood first issuehelp wanted
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
3 commentaires0 réaction0 personne assignée
Ouverte
[EVAL] Big-Bench Extra Hard (BBEH)
good first issuehelp wantednew-taskscience-team
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
3 commentaires0 réaction0 personne assignée
Ouverte
[EVAL] Add TUMLU benchmark
good first issuehelp wantednew-task
Pourquoi recommandéeAucune personne assignée · Label adapté aux débutants
Aucune personne assignéeLabel adapté aux débutants
10 commentaires0 réaction0 personne assignée