HSHKD A-Mark: An Audio Classification Model with Hard-Soft Hybrid Knowledge Distillation
摘要
Recently, the demand for real-time noise detection and classification has increased as a means to mitigate conflicts caused by inter-floor noise in residential buildings. However, in real home environments, various daily-life sounds often occur simultaneously and are mixed, and the label structures of public datasets do not clearly separate inter-floor noise from general everyday sounds. In this paper, we propose Mark, an audio classification model designed to distinguish noise types that are directly related to inter-floor noise from other noise. Mark consists of an ensemble of teacher models specialized for each noise type (Mark ensemble) and a single student model, and it is trained using a hard–soft hybrid knowledge distillation scheme that jointly exploits soft labels generated by the teacher ensemble and hard labels obtained from the collected data. Experiments on the AI-Hub inter-floor noise dataset show that the proposed model achieves balanced performance in both multi-class noise-type classification and binary classification of inter-floor versus non–inter-floor noise. In particular, for the binary task of distinguishing inter-floor noise from other noise, Mark attains an F1-score of approximately 0.91 with very high precision, demonstrating its effectiveness in detecting inter-floor noise in real residential environments while minimizing false alarms.