Pushing the boundaries of deepfake audio detection with a hybrid MFCC and spectral contrast approach
摘要
The proliferation of deepfake audio content presents a formidable challenge in today’s digital landscape, necessitating advanced detection techniques to combat misinformation and manipulation. In response, this study introduces a novel approach to deepfake audio detection, leveraging a unified framework that combines Mel-Frequency Cepstral Coefficients (MFCC) and spectral contrast features. The research commences by meticulously curating a comprehensive dataset of In-the-Wild Audio Data, ensuring a diverse representation of authentic and deepfake audio samples. Subsequently, the proposed framework extracts high-dimensional audio features from the dataset, utilizing both MFCC and spectral contrast analysis techniques. Through rigorous experimentation and model optimization, our approach demonstrates exceptional performance in distinguishing between genuine and deepfake audio recordings. Notably, the integration of MFCC DeltaDelta features with spectral contrast yields superior results, showcasing the efficacy of the combined feature set. Furthermore, the framework incorporates a streamlined Artificial Neural Network (ANN) architecture, optimizing computational efficiency without compromising detection accuracy. Empirical evaluation on benchmark datasets underscores the robustness and reliability of the proposed methodology, achieving an impressive accuracy score of 98.52%. Additionally, the introduction of the F1-score evaluation metric further reinforces the effectiveness of our approach. By providing a clear and systematic framework for deepfake audio detection, this research significantly enhances the capabilities of forensic investigators and digital media analysts in identifying and mitigating the proliferation of deceptive audio content. The proposed approach represents a crucial step towards combating the evolving threat of deepfake technology in the digital age.