错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Study of Various Audio Augmentation Methods and Their Impact on Automatic Speech Recognition

  • Naorem Karline Singh,
  • Yambem Jina Chanu,
  • Hoomexsun Pangsatabam

摘要

Automatic speech recognition transforms a spoken utterance into its corresponding textual form. The scarcity of annotated speech data hinders the development of any language’s ASR system. The quantity and quality of training data that are provided have a direct correlation with how well these systems perform. This is further excruciating for regional low-resource languages. The development of a credible speech corpus for such a language is time-consuming and expensive. To alleviate this problem to some extent, data augmentation is generally employed to enhance the generalization ability of the model. This paper examine several currently used techniques, ranging from raw data augmentation to frequency domain augmentation, such as time masking, peak normalization, and frequency masking. We analyze and contrast their benefits and drawbacks in terms of a few key factors such as execution time and computational accuracy. We also look at the applicability of a few techniques that have been shown to improve speech recognition through audio data augmentation.