facebookresearch/fairseq

reduce padding overhead when using buckets

Aperta

#5011 aperta il 6 mar 2023

 (1 commento) (1 reazione) (0 assegnatari)Python (6224 fork)batch import
enhancementhelp wantedneeds triage

Metriche repository

Star
 (29.107 stelle)
Metriche merge PR
 (Metriche PR in attesa)

Descrizione

🚀 Feature Request

The current get_buckets function ensures that you have the same number of samples in each bucket. This can result in many unnecessary padding. The feature would enable the user to set an optimal bucket size to reduce padding.

Motivation

padding = unrequited work -> bad performance. While finding the optimal bucket can take 1-5 minutes it can reduce training time by 10-50%.

Pitch

Add support for using Integer programming to set the optimal bucketing size. The implementation I had in mind depends on pulp. I can share the code for optimal bucketing.

Alternatives

k-means algorithm can give a suboptimal but faster approximation

Additional context

Guida contributor