unslothai/unsloth

DPO, ORPO - grad accumulation fix

Chiusa

#1178 aperta il 24 ott 2024

 (1 commento) (0 reazioni) (0 assegnatari)Python (5658 fork)batch import
feature requesthelp wanted

Metriche repository

Star
 (64.271 stelle)
Metriche merge PR
 (Merge medio 3g 15h) (525 PR mergiate in 30 g)

Descrizione

Goal: Propagate gradient accumulation fix to DPO - much harder since it requires a full rewrite of https://github.com/huggingface/trl/blob/main/trl/trainer/dpo_trainer.py

Guida contributor