kohya-ss/sd-scripts

Question about V-Prediction in SDXL Finetuning

開放

#1,163 建立於 2024年3月9日

 (1 則留言) (0 個反應) (0 位負責人)Python (1,217 個分叉)batch import
help wanted

倉庫指標

星標
 (7,201 顆星)
PR 合併指標
 (平均合併 18小時 39分鐘) (30 天內合併 16 個 PR)

描述

It's just my one-sided doubts, about the implement of the v-prediction. In sdxl training, the source code implements v-prediction by:

def add_v_prediction_like_loss(loss, timesteps, noise_scheduler, v_pred_like_loss):
    scale = get_snr_scale(timesteps, noise_scheduler)
    # print(f"add v-prediction like loss: {v_pred_like_loss}, scale: {scale}, loss: {loss}, time: {timesteps}")
    loss = loss + loss / scale * v_pred_like_loss
    return loss

which is mathematically equivalent to: L:=L+snr*L*w, where w=v_pred_like_loss, and snr=scale, while the paper suggests: L:=snr*L.

So, is the source adds additional v-pred like loss rather than scaling it? Why are the implementation and paper different? I'm not a mathematician, and maybe I'm short-sighted. Hope someone can answer my doubts :D

貢獻者指南