enhancementgood first issue
仓库指标
- 星标
- (23,049 个星标)
- PR 合并指标
- (平均合并 3天 6小时) (30 天内合并 125 个 PR)
描述
According to the documentation, veRL only supports AutoModelForSequenceClassification. What would be the best way to implement generative reward model (GenRM) for veRL? I tried looking at FSDP Workers and Megatron-LM Workers but they no longer exist.