facebookresearch/fairseq

reduce padding overhead when using buckets

オープン

#5,011 opened on 2023/03/06

 (1 件のコメント) (1 件のリアクション) (0 人の担当者)Python (6,224 件のフォーク)batch import
enhancementhelp wantedneeds triage

Repository metrics

Stars
 (29,107 個のスター)
PR merge metrics
 (PR metrics pending)

説明

🚀 Feature Request

The current get_buckets function ensures that you have the same number of samples in each bucket. This can result in many unnecessary padding. The feature would enable the user to set an optimal bucket size to reduce padding.

Motivation

padding = unrequited work -> bad performance. While finding the optimal bucket can take 1-5 minutes it can reduce training time by 10-50%.

Pitch

Add support for using Integer programming to set the optimal bucketing size. The implementation I had in mind depends on pulp. I can share the code for optimal bucketing.

Alternatives

k-means algorithm can give a suboptimal but faster approximation

Additional context

コントリビューターガイド