Whisper+AASIST for DeepFake Audio Detection
摘要
This study introduces a novel approach, combining the Whisper model with the AASIST architecture to enhance the detection performance of deepfake audio. Termed Whisper \(+\) AASIST, our investigation demonstrates its competitive edge over existing models on the 2021 ASVspoof DF subset. Notably, it surpasses these models in the challenging In-the-Wild datasets. This innovative fusion signifies a significant leap in deepfake audio detection, showcasing the effectiveness of synergistic model architectures. Through a comprehensive analysis, we unravel the potential of this hybrid approach in addressing evolving challenges within the domain of deceptive audio content. The findings underscore the significance of such inventive model combinations, providing a foundation for further advancements in deepfake audio detection methodologies.