Semi-supervised Learning for Audio Signal
摘要
Semi-supervised learning (Semi-SL) is a machine learning paradigm that combines labeled and unlabeled data to enhance model performance, has emerged as a powerful tool in audio signal processing–particularly in scenarios where labeled data is scarce or costly to acquire. This chapter explores the principles of Semi-SL applied to audio signals, highlighting various techniques such as pseudo-labeling, consistency regularization, and graph-based methods. Building on these foundations, this chapter presents typical applications of Semi-SL in tasks like speech recognition, sound event detection, and music classification, alongside an analysis of notable Semi-SL-based models driving advancements in the field. This chapter discuss the challenges and opportunities of Semi-SL in the audio domain, including issues of data sparsity, class imbalance, and domain adaptation. Moreover, We critically examine these limitations while highlighting opportunities to address them through innovative algorithmic and data-centric strategies. Finally, we synthesize emerging trends and open research questions, offering insights into the future potential of Semi-SL for advancing audio signal processing in both theoretical and practical contexts.