Self-distillation framework for improving fake speech detection in the domain variability scenario
摘要
Robust fake speech detection systems are crucial in an era where audio recordings can be easily altered or developed due to advancements in technology. The potential impact of this technology could be devastating due to its susceptibility to misuse. It can lead to political and economic instability, severe financial repercussions, the spread of misinformation, defamation, and security breaches in ASV systems and facilitate theft and fraud. Further, existing detection systems face a key issue; most state-of-the-art systems only provide results on the corpus they are trained on but fail in the domain variability scenario. The goal of this study is to enable a system to generalize across various domains, enhancing its reliability in real-world scenarios and thereby solidifying the authenticity of speech, making speech-based systems more dependable. All three audio spoofing scenarios are considered, logical access-based (LA) attacks including TTS and voice conversion (VC) techniques, spoofing attacks generated in real physical space, namely physical access (PA) attacks (replay, mimicry) and advanced deepfake technologies. The proposed system put forth in this study achieves commendable performance on three diverse test datasets. Notably, EER of 0.286, 0.337 and 0.371 is achieved on the In-The-Wild dataset with the proposed system implemented on ResNet, ECANet and SENet.