Hierarchical Discrepancy-Aware Interaction Network for Face Forgery Detection
摘要
Deep learning technologies, such as deepfakes, have garnered significant attention in both industry and academia, especially those employed for generating forged facial images. These automated facial manipulation technologies can replace the original facial images, even in videos, with any target object while preserving its facial behaviors, which potentially triggers serious security threats to society. Current face forgery detection methods follow a strong assumption to primarily concentrate on exploring specific forgery clues determined by feature distribution variations, such as local texture, frequency differences, and noise distortion. However, these learned tampering cues often exhibit unpredictability in terms of size, position, and quantity. Most methods tend to overemphasize the most conspicuous features, leading to insufficient consideration of all the acquired knowledge and thereby limiting robustness and generalization capabilities. In this paper, we propose a novel hierarchical discrepancy-aware interaction learning (HDIL) strategy to equally concern all acquired information for realizing multilevel facial concerns extraction. Specifically, we introduce a diversified semantic controller to generate different semantic components according to the proportion of valid information to enlarge feature mapping. Furthermore, we design a hierarchical discrepancy-aware interaction learning scheme to aggregate multiple forged components across different spaces as well as increase the spatial correlations by presenting a specific perceptual attention module. We conducted extensive experimental evaluations on widely used datasets to validate the effectiveness of the proposed method.