buggood first issuehelp wantedscience-team
倉庫指標
- 星標
- (2,496 顆星)
- PR 合併指標
- (PR 指標待抓取)
描述
Describe the bug
For now tokenization is bing made in a for loop, making the whole process very expensive for large benchmarks.
To Reproduce
run any big benchmarks
Expected behavior
Use batch tokenize to speed it up.
Version info
Please provide your operating system, lighteval version or commit if you installed from main, and pip/conda environment if your problem concerns dependencies.