unslothai/unsloth

DPO, ORPO - grad accumulation fix

クローズ

#1,178 opened on 2024/10/24

 (1 件のコメント) (0 件のリアクション) (0 人の担当者)Python (5,658 件のフォーク)batch import
feature requesthelp wanted

Repository metrics

Stars
 (64,271 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Goal: Propagate gradient accumulation fix to DPO - much harder since it requires a full rewrite of https://github.com/huggingface/trl/blob/main/trl/trainer/dpo_trainer.py

コントリビューターガイド