Comparative Analysis of Pretrained Models for Speech Enhancement in Noisy Environments
摘要
Speech Enhancement is the set of techniques and algorithms aimed at enhancing the overall quality of speech signals across diverse conditions both qualitatively and quantitatively. Speech enhancement aims to enhance voice signals whose quality has been diminished by various kinds of noise or distortion. Different techniques were adopted in previous years. Researchers have started working with Machine Learning techniques recently, prior to which they have followed traditional methods like Wiener Filtering, Spectral Subtraction, etc. The advancement of machine learning techniques day by day has laid the path for our work. Our work is to investigate the performance of three models viz., ESPNet-SE, SpeechBrain MetricGAN+ and SpeechBrain SepFormer models on a mixed dataset namely VoiceBank and Demand, which has added noise on clean signals. Among all the models, SpeechBrain MetricGAN+ performed well by approximately 30.05% on ESPNet-SE and 10.29% on SpeechBrain SepFormer models. Trained models are publicly available.