错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatial and Frequency Domains Inconsistency Learning for Face Forgery Detection

  • Caili Gao,
  • Peng Qiao,
  • Yong Dou,
  • Qisheng Xu,
  • Xifu Qian,
  • Wenyu Li

摘要

With the rapid development of face forgery technology, it has attracted widespread attention. The current face forgery detection methods, whether based on Convolutional Neural Network (CNN) or Vision Transformer (ViT), are biased towards extracting local or global features respectively. These methods are relatively one-sided, and the extracted features are not robust and general enough. In this work, we exploit intra-frame inconsistency as well as inter-modal inconsistency between spatial and frequency domains to improve performance and generalization for face forgery detection. We efficiently extract intra-frame inconsistency by utilizing the capabilities of Swin Transformer. Its self-attention mechanism and attention mapping between patch embeddings naturally represent the inconsistency relations, allowing for simultaneous modeling of both local and global features, making it our ideal choice. Meanwhile, we also introduce frequency information to further improve detection performance, and design a Cross-Attention Feature Fusion (CAFF) module to exploit the inconsistency between spatial and frequency modalities to extract more general feature representations. Extensive experiments demonstrate the effectiveness of the proposed method.