ggml-org/llama.cpp
Study how LM Evaluation Harness works and try to implement it
开放
#231 创建于 2023年3月17日
enhancementgeneration qualityhelp wantedhigh priorityresearch 🔬
仓库指标
- 星标
- (125,444 个星标)
- PR 合并指标
- (PR 指标待抓取)
描述
Update 10 Apr 2024: https://github.com/ggerganov/llama.cpp/issues/231#issuecomment-2047759312
It would be great to start doing this kind of quantitative analysis of ggml-based inference:
https://bellard.org/ts_server/
It looks like Fabrice evaluates the models using something called LM Evaluation Harness:
https://github.com/EleutherAI/lm-evaluation-harness
I have no idea what this is yet, but would be nice to study it and try to integrate it here and in other ggml-based projects.
This will be very important step needed to estimate the quality of the generated output and see if we are on the right track.