kohya-ss/sd-scripts

Question about V-Prediction in SDXL Finetuning

开放

#1,163 创建于 2024年3月9日

 (1 条评论) (0 个反应) (0 位负责人)Python (1,218 个派生)batch import
help wanted

仓库指标

星标
 (7,198 个星标)
PR 合并指标
 (平均合并 18小时 39分钟) (30 天内合并 16 个 PR)

描述

It's just my one-sided doubts, about the implement of the v-prediction. In sdxl training, the source code implements v-prediction by:

def add_v_prediction_like_loss(loss, timesteps, noise_scheduler, v_pred_like_loss):
    scale = get_snr_scale(timesteps, noise_scheduler)
    # print(f"add v-prediction like loss: {v_pred_like_loss}, scale: {scale}, loss: {loss}, time: {timesteps}")
    loss = loss + loss / scale * v_pred_like_loss
    return loss

which is mathematically equivalent to: L:=L+snr*L*w, where w=v_pred_like_loss, and snr=scale, while the paper suggests: L:=snr*L.

So, is the source adds additional v-pred like loss rather than scaling it? Why are the implementation and paper different? I'm not a mathematician, and maybe I'm short-sighted. Hope someone can answer my doubts :D

贡献者指南