facebookresearch/fairseq

Speech Data Augmentation

Aberta

#4.971 aberto em 3 de fev. de 2023

 (0 comentário) (0 reação) (0 responsável)Python (6.224 forks)batch import
enhancementhelp wanted

Métricas do repositório

Stars
 (29.107 estrelas)
Métricas de merge de PR
 (Métricas PR pendentes)

Description

Note: this is issue is part of MLH fellowship

We would like to have more speech data augmentation in Fairseq.

I've been using augly which uses a lot of torchaudio.sox_effects, but I haven't been quite happy with the performance. All the sox_effects happen on CPU, which can be slow. I think some of them would better be implemented on GPU.

torchaudio tutorial explain how to implement some data augmentation, and their samples would work on GPU.

https://pytorch.org/audio/stable/tutorials/audio_data_augmentation_tutorial.html

I'm interested in reverb, change_volume, add_background_noise, tempo and pitch_shift.

Guia do colaborador