Speech Enhancement: Data Manipulation Techniques for Augmenting Existing Datasets
摘要
In everyday life, we frequently encounter challenges in voice communication due to noise interference, such as when using cell phones in cars or trains, in busy roadside areas, or markets. Speech enhancement techniques aim to extract clear speech signals from noisy backgrounds. Despite the recent advances in the use of Deep Neural Networks (DNNs), speech enhancement techniques, there are still challenges in adapting neural networks to handle unforeseen noise types, which can lead to reduced model performance. This paper provides a review of current techniques used in speech enhancement tasks and assesses the performance of BSRNN model across various noise types using the scheme of the On-the-fly Data mixing technique. The objective is to assess the effectiveness of the data augmentation method and its ability to adapt to different noise scenarios. By identifying limitations and strengths, this experimental research contributes to the advancement of speech enhancement technology and its applicability in real-world noisy environments.