facebookresearch/fairseq

Speech Data Augmentation

Ouverte

#4 971 ouverte le 3 févr. 2023

 (0 commentaire) (0 réaction) (0 personne assignée)Python (6 224 forks)batch import
enhancementhelp wanted

Métriques du dépôt

Stars
 (29 107 étoiles)
Métriques de merge PR
 (Métriques PR en attente)

Description

Note: this is issue is part of MLH fellowship

We would like to have more speech data augmentation in Fairseq.

I've been using augly which uses a lot of torchaudio.sox_effects, but I haven't been quite happy with the performance. All the sox_effects happen on CPU, which can be slow. I think some of them would better be implemented on GPU.

torchaudio tutorial explain how to implement some data augmentation, and their samples would work on GPU.

https://pytorch.org/audio/stable/tutorials/audio_data_augmentation_tutorial.html

I'm interested in reverb, change_volume, add_background_noise, tempo and pitch_shift.

Guide contributeur