Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Inquiry about the Self-Taught Evaluator

Open
#22 0 comments 0 reactions 2 assignees View on GitHub

@xianxl is already working on this.

Since Jan 8, 2025.

Assessment

This issue has not been assessed yet.

Description

The paper mentions that over 20,000 instructions categorized as “reasoning” were ultimately selected. However, the number of “reasoning”-related instructions obtained from WildChat far exceeds 20,000. How can the final 20,000 instructions used in the study be selected from this larger set? Following the settings described in the paper, I trained the model for 2 epochs, but my experiments indicate that the peak performance on RewardBench appears to be achieved after about one epoch, after which it begins to decline.

Dominant language
Python
Stars
382
Forks
49
PR merge metrics
No merged PRs in 30d

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from facebookresearch/RAM

All issues in facebookresearch/RAM

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.