Audio Classification
摘要
Audio classification constitutes a fundamental task in audio signal processing, aiming to categorize sounds into predefined classes based on their acoustic characteristics. Informed by the signal representations and machine learning techniques presented in Part II, this chapter examines theoretical foundations and practical implementations of audio classification systems. After introducing the general architecture of audio classification systems, the chapter focuses on three key tasks with varying granularity and temporal precision: acoustic scene classification for identifying environments, audio tagging for recognizing sound categories, and sound event detection for localizing sound events in time. For each task, the chapter provides a comprehensive review of developments, ranging from historical evolution to state-of-the-art methods, encompassing datasets, model architectures, learning paradigms, evaluation metrics, and recent advances. Through this systematic examination, the chapter offers valuable insights into the challenges, methodologies, and future directions in audio classification research and applications.