7 comments (7 comments)0 reactions (0 reactions)1 assignee (1 assignee)Python514 forks (514 forks)auto 404
good first issuehelp wantednew-task
Repository metrics
- Stars
- 2,496 stars (2,496 stars)
- PR merge metrics
- PR metrics pending (PR metrics pending)
Description
Evaluation short description
Evaluation metadata
Provide all available
Contributor guide
- Research direction
- Implement the Long Horizon Execution evaluation task using the provided paper and dataset. Study the paper (https://arxiv.org/abs/2509.09677) for metrics and methodology, and integrate the dataset from Hugging Face. Reference existing evaluation tasks in lighteval as templates.
- Tech stack
- python
- Domain
- machine learningai
- Issue type
- Feature
- Prerequisites
- Pythonlighteval framework