DFST: Dual-Stream Frequency-Spatial Transformer for Robust Deepfake Detection
摘要
With the rapid advancement of deepfake technologies, detection methods face challenges due to the diversity of forged samples and adversarial attacks. We propose a dual-stream frequency-spatial fusion Transformer (DFST) framework, comprising a spatiotemporal feature branch based on Vision Transformer to capture facial structure and dynamic semantic cues, and an adaptive frequency fingerprint branch that uses learnable multi-scale spectral filters to extract deepfake traces in the frequency domain. The two branches interact through a cross-modal co-attention mechanism, and an entropy-based adaptive weighting strategy balances their contributions to improve generalization. A perturbation sensitivity module is introduced to enhance robustness against adversarial attacks. Evaluated on the FaceForensics++ dataset under cross-manipulation settings, DFST achieves 90.02% average AUC with approximately 20.3M parameters, demonstrating strong accuracy and generalization. Results indicate that frequency-domain features provide complementary information for deepfake detection, and the fusion framework maintains stability across diverse forgery types and adversarial scenarios.