confident-ai/deepeval

Support popular benchmarks

开放

#508 创建于 2024年2月22日

 (11 条评论) (0 个反应) (0 位负责人)Python (1,677 个派生)auto 404
enhancementhelp wanted

仓库指标

星标
 (16,939 个星标)
PR 合并指标
 (平均合并 4天 16小时) (30 天内合并 26 个 PR)

描述

Currently, users of deepeval can only create their own evaluation dataset/test cases. To support more users fine-tuning their model, deepeval should be able to import standard benchmarks such as MMLU, hellaswag, TruthfulQA, Big Bench, etc.

Comment or DM on discord to discuss more: https://discord.com/invite/a3K9c8GRGt

贡献者指南