verl-project/verl

[FR] Documentation about Resolving OOM

已關閉

#1,014 建立於 2025年4月10日

 (6 則留言) (7 個反應) (0 位負責人)Python (4,440 個分叉)auto 404
call for contributiondocumentationgood first issue

倉庫指標

星標
 (23,087 顆星)
PR 合併指標
 (平均合併 3天 6小時) (30 天內合併 125 個 PR)

描述

Motivation

There are many issues related to OOM, e.g. #328 . We might need a clear guide about how to resolve OOM.

Plan

A non-exclusive enumeration about related configurations:

  1. Rollout:gpu_memory_utilization
  2. Other Inference:
    1. Liger Kernel
    2. *_max_len_per_gpu / micro_batch_size_per_gpu
  3. Training:
    1. Liger Kernel
    2. Ulysses Sequence Parallelism
    3. gradient checkpointing
    4. offload

TODO

  • Complete the list of related configurations
  • Benchmark the effect & overhead of each configuration

貢獻者指南