A Study of Various Audio Augmentation Methods and Their Impact on Automatic Speech Recognition
摘要
Automatic speech recognition transforms a spoken utterance into its corresponding textual form. The scarcity of annotated speech data hinders the development of any language’s ASR system. The quantity and quality of training data that are provided have a direct correlation with how well these systems perform. This is further excruciating for regional low-resource languages. The development of a credible speech corpus for such a language is time-consuming and expensive. To alleviate this problem to some extent, data augmentation is generally employed to enhance the generalization ability of the model. This paper examine several currently used techniques, ranging from raw data augmentation to frequency domain augmentation, such as time masking, peak normalization, and frequency masking. We analyze and contrast their benefits and drawbacks in terms of a few key factors such as execution time and computational accuracy. We also look at the applicability of a few techniques that have been shown to improve speech recognition through audio data augmentation.