confident-ai/deepeval

Support popular benchmarks

Aperta

#508 aperta il 22 feb 2024

 (11 commenti) (0 reazioni) (0 assegnatari)Python (1677 fork)auto 404
enhancementhelp wanted

Metriche repository

Star
 (16.939 stelle)
Metriche merge PR
 (Metriche PR in attesa)

Descrizione

Currently, users of deepeval can only create their own evaluation dataset/test cases. To support more users fine-tuning their model, deepeval should be able to import standard benchmarks such as MMLU, hellaswag, TruthfulQA, Big Bench, etc.

Comment or DM on discord to discuss more: https://discord.com/invite/a3K9c8GRGt

Guida contributor