Deepfake audio technology, while beneficial for entertainment and assistive technology, poses significant threats to societal trust and security, including disseminating misinformation and impersonation risks. To address these challenges, this paper proposes a novel approach that combines Sinc and wavelet filters in the feature extraction layer of the RawNet2 architecture. Experimental results demonstrate that the networks trained with these combined filters show better generalization capabilities across diverse datasets than the architecture of the original RawNet2. They repeatedly achieve superior results in detecting deepfake audio, demonstrating their ability to effectively identify subtle characteristics in spoofed audio signals. However, further research and refinement are necessary for optimal performance across all datasets. Future research directions include investigating optimal combinations of wavelets for deepfake audio detection and exploring other variations or combinations of filters to improve detection accuracy. Overall, the proposed approach offers a promising solution for detecting and mitigating the risks associated with deepfake audio, contributing to developing more robust and reliable deepfake detection systems for real-world applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deepfake Audio Detection with Sinc and Wavelet Filters in RawNet2

  • Uliana Zbezhkhovska,
  • Oleksandr Khapilin

摘要

Deepfake audio technology, while beneficial for entertainment and assistive technology, poses significant threats to societal trust and security, including disseminating misinformation and impersonation risks. To address these challenges, this paper proposes a novel approach that combines Sinc and wavelet filters in the feature extraction layer of the RawNet2 architecture. Experimental results demonstrate that the networks trained with these combined filters show better generalization capabilities across diverse datasets than the architecture of the original RawNet2. They repeatedly achieve superior results in detecting deepfake audio, demonstrating their ability to effectively identify subtle characteristics in spoofed audio signals. However, further research and refinement are necessary for optimal performance across all datasets. Future research directions include investigating optimal combinations of wavelets for deepfake audio detection and exploring other variations or combinations of filters to improve detection accuracy. Overall, the proposed approach offers a promising solution for detecting and mitigating the risks associated with deepfake audio, contributing to developing more robust and reliable deepfake detection systems for real-world applications.