Detecting Audio Deepfakes Through Emotional Fingerprinting
摘要
This study proposes a method to detect audio deepfakes by leveraging the inherent difficulty of replicating human emotions using deep learning models. We introduce a detection framework that utilizes Valence-Arousal-Dominance (VAD), which estimates emotional states in speech and extracts auxiliary features. These features, combined with latent embeddings from an efficient speech representation model, significantly improve detection accuracy. Our experimental evaluation demonstrates that the proposed approach enhances state-of-the-art methods like AASIST across multiple languages and environmental settings. This research suggests that incorporating emotional cues could be a promising approach to tackling the evolving challenge of audio deepfakes.