Deepfake Video Detection Guided by Identity and Temporal Inconsistency
摘要
With the rapid development of deepfake technology, the need for detecting deepfake videos has become increasingly important. Current deepfake detection methods mainly rely on specific low-level texture clues present in deepfake videos. However, as the generation technology continues to improve, deepfakes are becoming more and more realistic, with fewer low-level artifacts, making detection more challenging. To address this challenge, we introduce the Identity Comparison Network (ICN), which leverages high-level semantic information to identify deepfakes by analyzing facial ID inconsistencies. The Spatial Comparison Network (SCN) extracts traditional image artifact information to enhance the detection process. Additionally, we introduce the Frames Comparison Network (FCN) to identify inconsistencies between frames within a video, leveraging the weaknesses inherent in frame-by-frame forgeries. We propose a new method that fuses the above information to detect forgeries at both the intra-frame and inter-frame level. Through experimental evaluation on FF++ and Celeb-DF, we have demonstrated the effectiveness and generalization ability of our method in detecting deepfake videos.