Deep fake Video Face Recognition Using Supervised Contrastive Learning for Scalability and Interpretability
摘要
An efficient and effective video forgery detector is essential to detect false facial images in videos to properly organize the footage and retrieve relevant information. Disseminating a fabricated (false) image containing illegal content through social media might have serious repercussions. Previous approaches mainly frame face forgery detection in the video as a classification challenge based on cross-entropy loss, rather than the underlying distinctions between genuine and fraudulent faces. As deepfakes continue to circulate, reliable methods that can identify deceptively realistic fakes are in high demand. Deepfake might be challenging to gather large amounts of data when required in many practical contexts. The research provides a cutting-edge technique for detecting deepfake videos using a supervised contrastive learning framework for local and global visual models (SCL-LGVM) to learn audio/video portrayals that generalize the tasks that involve both global contextual features for classification and local fine-grained–spatial knowledge. At first, generate two transformed versions of a video image and feed them into two sequential subnetworks, i.e., an encoder and a decoder. As a result of optimizing for two contradictory aims, a model can draw in both global and local image perception from acoustic inputs. The last step in supervised training involves optimizing the classifier's outputs to achieve the highest possible degree of correspondence. Extensive experiments show that the managed learning technique yields detection performance comparable to state-of-the-art unsupervised methods within and across datasets. A high-quality deepfake dataset, deepfake and real images [37] consisting of 4,000 deepfake videos were made using state-of-the-art face facial forgery techniques to encourage further study into deep phony detection. The suggested model reaches state-of-the-art performance, and thorough tests and analyses show the technique's resilience and generalizability.