confident-ai/deepeval

Support popular benchmarks

オープン

#508 opened on 2024/02/22

 (11 件のコメント) (0 件のリアクション) (0 人の担当者)Python (1,677 件のフォーク)auto 404
enhancementhelp wanted

Repository metrics

Stars
 (16,939 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Currently, users of deepeval can only create their own evaluation dataset/test cases. To support more users fine-tuning their model, deepeval should be able to import standard benchmarks such as MMLU, hellaswag, TruthfulQA, Big Bench, etc.

Comment or DM on discord to discuss more: https://discord.com/invite/a3K9c8GRGt

コントリビューターガイド