A Study on Domain Adaptation for Audio-Visual Speech Enhancement
摘要
This paper presents the DA-AVSE system developed for the ASRU 2023 Audio-Visual Speech Enhancement (AVSE) Challenge. We initially employed three well-established AVSE models: MEASE, MTMEASE, and PLMEASE. These models demonstrated effectiveness even without utilizing matched data for training. To further enhance the performance, we introduced a domain adaptation method. More specifically, we utilized pseudo-labels generated by the models above in conjunction with the official baseline to fine-tune each model. Through extensive experiments, we observed that our method significantly improved the models’ generalization to the target test set, regardless of whether the training and testing conditions matched. Additionally, we implemented a multi-model fusion strategy to enhance the overall model performance further. Our system exhibited significant improvements in all objective metrics, including PESQ, STOI, and SiSDR, compared to almost all competing teams. As a result, our system ranked the 2nd place in the objective metrics comparison for track 1.