ggml-org/llama.cpp

Study how LM Evaluation Harness works and try to implement it

Offen

#231 geöffnet am 17.03.2023

 (10 Kommentare) (9 Reaktionen) (0 zugewiesene Personen)C++ (22.115 Forks)batch import
enhancementgeneration qualityhelp wantedhigh priorityresearch 🔬

Repository-Metriken

Stars
 (125.444 Sterne)
PR-Merge-Metriken
 (PR-Metriken ausstehend)

Beschreibung

Update 10 Apr 2024: https://github.com/ggerganov/llama.cpp/issues/231#issuecomment-2047759312


It would be great to start doing this kind of quantitative analysis of ggml-based inference:

https://bellard.org/ts_server/

It looks like Fabrice evaluates the models using something called LM Evaluation Harness:

https://github.com/EleutherAI/lm-evaluation-harness

I have no idea what this is yet, but would be nice to study it and try to integrate it here and in other ggml-based projects. This will be very important step needed to estimate the quality of the generated output and see if we are on the right track.

Contributor Guide