Audio Denoising
摘要
In daily life, audio acquisition systems inevitably capture noise and interference along with the desired audio signals, which degrades audio quality and affects backend systems. To address this issue, audio denoising techniques have been proposed to remove noise or interference from noisy audio while preserving the target signals, thereby improving audio quality. Speech, as a typical case in audio denoising tasks, aims to retain the target speaker’s voice, eliminate interferences, and enhance speech quality/intelligibility, thereby boosting the performance of speech-related systems, e.g., voice communication, automatic speech recognition. Having been studied for decades, speech denoising has provided foundational ideas and methods for other audio denoising tasks. This chapter takes speech denoising as the core example to summarize signal models, denoising approaches, and particularly deep learning-based methods with typical architectures of audio denoising. Additionally, while previous literature rarely focuses on real-time audio denoising, deep learning-based real-time techniques have gained growing attention in recent years. Thus, this chapter reviews lightweight real-time denoising methods from perspectives of feature engineering, model design, etc. Finally, we outline the current challenges and future trends in audio denoising based on our observations.