BiFAT: Bilateral Filtering and Attention Mechanisms in a Two-Stream Model for Deepfake Detection
摘要
Addressing the significant societal concerns triggered by the widespread dissemination of Deepfake facial forgeries on the internet, and the shortcomings of existing Deepfake video detection methods in terms of generalizability and resistance to compression artifacts, we introduce a model named BiFAT. BiFAT synergizes bilateral filtering and attention mechanisms within a two-stream model to transcend the limitations of traditional binary classification approaches in Deepfake detection. Firstly, we employ an attention mechanism combined with the Steganalysis Rich Filters and Attention (SRMA) for spatial feature extraction, capturing intricate local textures and structures. Secondly, discrete Fourier transform and a complex adaptive filter are applied for frequency domain feature extraction, ensuring a comprehensive analysis of the image. This dual-domain approach, augmented by attention layers, refines the feature extraction and amalgamation process, significantly enhancing detection performance on benchmark datasets such as DF-1.0, DFDC, Celeb-DF, and FaceForensics++. Finally, our method demonstrates a notable improvement in model convergence speed, addressing the challenge of managing an excessive number of features, a common issue in contemporary DeepFake detection models, thereby boosting the model’s generalizability and compression artifact resistance.