A Domain Adaptive Dual-Teacher Network for Sound Event Detection
摘要
In the absence of abundant strongly labeled real data, traditional sound event detection faces the issue of distribution inconsistency between synthetic and real audio data. To address this, this paper proposes an intermediate domain generation strategy based on mixing feature statistics. By mixing features from the source and target domains, a transitional domain is created to alleviate data distribution differences and enhance the model's generalization capability. Additionally, this paper introduces an innovative dual-teacher network architecture. By leveraging high-quality pseudo-labels generated by expert networks, it improves the utilization efficiency of unlabeled data, thereby optimizing the performance of the teacher-student network. Experiments conducted on the DCASE Task4 dataset validate the effectiveness of the proposed method. The results show that the proposed DADT-Net significantly enhances detection performance in cross-domain sound event detection tasks, achieving a 12.5% improvement in the PSDS evaluation metric compared to traditional CRNN methods.