verl-project/verl

Additional memory optimization features

Aperta

#144 aperta il 27 gen 2025

 (6 commenti) (3 reazioni) (0 assegnatari)Python (4426 fork)auto 404
call for contributionenhancementgood first issue

Metriche repository

Star
 (23.023 stelle)
Metriche merge PR
 (Merge medio 3g 6h) (125 PR mergiate in 30 g)

Descrizione

  • Activation offloading (see implementation here)
  • Fusing optimizer step into backward pass (see implementation here)
  • Utilize full_shard reshard_after_forward (see here). I wasn't 100% sure if I could see this already implemented in veRL.

These optimizations largely trade off decreased peak memory useage for additional compute, so may only be useful for training larger models, and in GPU-constrained settings.

Guida contributor