facebookresearch/fairseq

Speech Data Augmentation

オープン

#4,971 opened on 2023/02/03

 (0 件のコメント) (0 件のリアクション) (0 人の担当者)Python (6,224 件のフォーク)batch import
enhancementhelp wanted

Repository metrics

Stars
 (29,107 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Note: this is issue is part of MLH fellowship

We would like to have more speech data augmentation in Fairseq.

I've been using augly which uses a lot of torchaudio.sox_effects, but I haven't been quite happy with the performance. All the sox_effects happen on CPU, which can be slow. I think some of them would better be implemented on GPU.

torchaudio tutorial explain how to implement some data augmentation, and their samples would work on GPU.

https://pytorch.org/audio/stable/tutorials/audio_data_augmentation_tutorial.html

I'm interested in reverb, change_volume, add_background_noise, tempo and pitch_shift.

コントリビューターガイド