Wrap evaluation benchmark using HF-trainer
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Refactor
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- huggingface, python
- Domain
- machine-learning
Research direction
Start by reviewing the repository's current benchmark evaluation flow and identifying the trainer entry points; the issue names no files or tests. Compare the existing task-specific handling with Hugging Face Trainer requirements for data_loader, DataCollator, compute_metrics, and optional predictions. Done means the benchmark is wrapped with Hugging Face Trainer while preserving task evaluation and supporting future fine-tuning.
Written by the indexing model from the issue text.
Description
This might sounds like a bit of re-structuring but for the sake of future compatibility, I propose the following,
- Move to
huggingfacetrainer: This will help the repo to automatically adapt todeepspeedand all the exclusive features of transformers library. - We don't have to re-invent the wheel. Given that we are using huggingface trainer, we only need to implement the following functions for a trainer for different tasks.
--data_loader
--DataCollator
--compute_metrics
--predictions(if needed) - In case if we want to
finetuneour full model, we don't have to change a lot in the surface level.
I would love to take some responsibility if needed. Let me know. @jaketae @tianjianjiang @wilsonyhlee
- Dominant language
- Python
- Stars
- 42
- Forks
- 24
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from bigscience-workshop/evaluation
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
-
Refactor task template to merge multilingual.json and english.jsonMay be free again @wilsonyhlee claimed this 1855 days ago, and no pull request is open. Open
bigscience-workshop/evaluation#64 · 1 assignee ·
-
Setup testingMay be free again @tttyuntian claimed this 1870 days ago, and no pull request is open. Openengineering
bigscience-workshop/evaluation#57 · 2 comments · 3 assignees ·
-
Start overleaf for benchmark tech reportMay be free again @arunraja-hub claimed this 1866 days ago, and no pull request is open. Opendocumentation
bigscience-workshop/evaluation#54 · 1 comment · 3 reactions · 2 assignees ·
-
translate validation prompts into all training languagesMay be free again @epavlick claimed this 1877 days ago, and no pull request is open. Openmultilingual simple_benchmark
bigscience-workshop/evaluation#50 · 9 comments · 2 assignees ·
All issues in bigscience-workshop/evaluation
Similar issues
-
New InternshipOpennew_internship
Difficulty 1/5 Under an hour Newbie friendliness 70/100
-
[BUG] Reports tab: "Unban" button tooltip shows raw `{{ip}}` placeholder instead of the IP addressOpenbug javascript ui
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
bunkerity/bunkerweb#4001 · 1 comment ·
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 92/100
PedestrianDynamics/pyFDS-Evac#476 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
google/differential-privacy#516 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
adobe-fonts/source-serif#153 ·