Inquiry about the Self-Taught Evaluator
@xianxl is already working on this.
Since Jan 8, 2025.
Assessment
This issue has not been assessed yet.
Description
The paper mentions that over 20,000 instructions categorized as “reasoning” were ultimately selected. However, the number of “reasoning”-related instructions obtained from WildChat far exceeds 20,000. How can the final 20,000 instructions used in the study be selected from this larger set? Following the settings described in the paper, I trained the model for 2 epochs, but my experiments indicate that the peak performance on RewardBench appears to be achieved after about one epoch, after which it begins to decline.
- Dominant language
- Python
- Stars
- 382
- Forks
- 49
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from facebookresearch/RAM
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
facebookresearch/RAM#77 ·
-
facebookresearch/RAM#13 · 1 comment · 1 assignee ·
All issues in facebookresearch/RAM
Similar issues
-
bug confirmed issue
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
open-webui/open-webui#30750 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
OpenwaterHealth/openmotion-bloodflow-app#604 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
good first issue
Difficulty 1/5 Under an hour Newbie friendliness 90/100