Active Learning for Audio Signal
摘要
This chapter examines how Active Learning (AL) can be applied to audio signal processing to tackle the challenge of limited and expensive labeled data. Rather than labeling large datasets, AL improves efficiency by strategically selecting the most informative samples for annotation. We first outline the general AL framework, then explore different query strategies, and how they can be used in tasks like sound event classification, automatic speech recognition (ASR), and emotion recognition. We also discuss pool-based sampling, which helps reduce manual labeling efforts. Finally, we highlight the benefits and challenges of integrating AL into deep learning (DL) for audio applications, showing how it can boost model performance while lowering annotation costs.