Audio Source Separation
摘要
Audio source separation (ASS) is a core task in computational auditory scene analysis, aiming to isolate individual sound sources from complex mixtures. This mirrors the remarkable human ability to focus on a specific voice or sound in noisy environments-a phenomenon often described as the “cocktail party effect.” Due to its wide range of applications in speech, music, and environmental sound processing, ASS has become a critical research area in audio signal processing. This chapter introduces the scope and applications of ASS, outlines its mathematical foundations, and discusses signal models and evaluation metrics. We then delve into three key domains: speech separation, music source separation, and audio event source separation. Each section traces the evolution of methodologies while showcasing state-of-the-art approaches that have dramatically improved separation quality across diverse acoustic environments. The techniques presented in this chapter not only advance academic understanding of audio processing but also enable critical real-world applications in speech enhancement, hearing aids, music production, audio forensics, and human-computer interaction, providing readers with essential knowledge to address increasingly complex audio separation challenges in both research and industry settings.