verl-project/verl

Additional memory optimization features

Ouverte

#144 ouverte le 27 janv. 2025

 (6 commentaires) (3 réactions) (0 personne assignée)Python (4 440 forks)auto 404
call for contributionenhancementgood first issue

Métriques du dépôt

Stars
 (23 087 étoiles)
Métriques de merge PR
 (Merge moyen 3j 6h) (125 PRs mergées en 30 j)

Description

  • Activation offloading (see implementation here)
  • Fusing optimizer step into backward pass (see implementation here)
  • Utilize full_shard reshard_after_forward (see here). I wasn't 100% sure if I could see this already implemented in veRL.

These optimizations largely trade off decreased peak memory useage for additional compute, so may only be useful for training larger models, and in GPU-constrained settings.

Guide contributeur