This paper focuses on the creation of a model for distinguishing clean speech (speech without music, interaction or background sounds) from speech containing noise. This model can be used to exploit several speech recognition models such as: automatic emotion recognition models, speaker recognition models, language recognition models, etc.. These models are usually trained on data sets containing short-length files. To be useful, these models need to be used in real-world environments where we use voice recordings of several minutes or hours in length which often contain noise. Keeping these noise will lead to false predictions from the model. It is therefore important to be able to automatically distinguish pure speech files from noisy files in order to provide the model with good files. Our contribution in this paper is twofold. Firstly, we are making available to the scientific community a dataset of around 15.21 h of speech containing clean utterances and noisy utterances. Three languages are considered in this paper: Moore, Dioula and Fulfulde. The second contribution concerns the pure speech detection model, which is useful for exploiting other speech recognition models. The combination of MFCC+F0 extractor is used to have speech characteristics and the LSTM is applied as a deep learning algorithm to classify audio into “noise” and “no noise” classes. Evaluation of this model gives an accuracy of 97.62%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Pure Speech Files Detection Using MFCC+F0 Features and LSTM Classifier in Multilingual Context for Low-Resource Languages

  • Go Issa Traore,
  • Borlli Michel Jonas Some

摘要

This paper focuses on the creation of a model for distinguishing clean speech (speech without music, interaction or background sounds) from speech containing noise. This model can be used to exploit several speech recognition models such as: automatic emotion recognition models, speaker recognition models, language recognition models, etc.. These models are usually trained on data sets containing short-length files. To be useful, these models need to be used in real-world environments where we use voice recordings of several minutes or hours in length which often contain noise. Keeping these noise will lead to false predictions from the model. It is therefore important to be able to automatically distinguish pure speech files from noisy files in order to provide the model with good files. Our contribution in this paper is twofold. Firstly, we are making available to the scientific community a dataset of around 15.21 h of speech containing clean utterances and noisy utterances. Three languages are considered in this paper: Moore, Dioula and Fulfulde. The second contribution concerns the pure speech detection model, which is useful for exploiting other speech recognition models. The combination of MFCC+F0 extractor is used to have speech characteristics and the LSTM is applied as a deep learning algorithm to classify audio into “noise” and “no noise” classes. Evaluation of this model gives an accuracy of 97.62%.