facebookresearch/fairseq

reduce padding overhead when using buckets

开放

#5,011 创建于 2023年3月6日

 (1 条评论) (1 个反应) (0 位负责人)Python (6,224 个派生)batch import
enhancementhelp wantedneeds triage

仓库指标

星标
 (29,107 个星标)
PR 合并指标
 (PR 指标待抓取)

描述

🚀 Feature Request

The current get_buckets function ensures that you have the same number of samples in each bucket. This can result in many unnecessary padding. The feature would enable the user to set an optimal bucket size to reduce padding.

Motivation

padding = unrequited work -> bad performance. While finding the optimal bucket can take 1-5 minutes it can reduce training time by 10-50%.

Pitch

Add support for using Integer programming to set the optimal bucketing size. The implementation I had in mind depends on pulp. I can share the code for optimal bucketing.

Alternatives

k-means algorithm can give a suboptimal but faster approximation

Additional context

贡献者指南