Our research aims to enhance the accuracy of Amazigh speech recognition through a data augmentation method. We employ Convolutional Neural Networks (CNNs) for speech recognition, utilizing Mel Spectrograms extracted from audio files. This study focuses on recognizing the first ten Amazigh digits. Our approach centers on refining recognition accuracy and augmenting data by modifying the number of bands in the filter bank. Additionally, our study involves 42 speakers who have participated, applying a speaker-independent approach throughout the experiments. We conducted three experiments using original data and three additional experiments using a combination of original and augmented data. The augmentation method involved changing the number of bands in the filter bank. Through these experiments with different CNN models, one model exhibited a 2.89% increase in accuracy, contributing to the overall improvement in speech recognition accuracy, with the highest achieved level reaching 95.88%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Augmentation for Amazigh Speech Recognition Using Filter Banks

  • Hossam Boulal,
  • Mohamed Hamidi,
  • Jamal Barkani,
  • Mustapha Abarkan

摘要

Our research aims to enhance the accuracy of Amazigh speech recognition through a data augmentation method. We employ Convolutional Neural Networks (CNNs) for speech recognition, utilizing Mel Spectrograms extracted from audio files. This study focuses on recognizing the first ten Amazigh digits. Our approach centers on refining recognition accuracy and augmenting data by modifying the number of bands in the filter bank. Additionally, our study involves 42 speakers who have participated, applying a speaker-independent approach throughout the experiments. We conducted three experiments using original data and three additional experiments using a combination of original and augmented data. The augmentation method involved changing the number of bands in the filter bank. Through these experiments with different CNN models, one model exhibited a 2.89% increase in accuracy, contributing to the overall improvement in speech recognition accuracy, with the highest achieved level reaching 95.88%.